LoRA
Nvidia's router dealt the cards evenly. That was the whole problem.
Nvidia PAIR split 20 requests 10/10 across an RTX 3090 and a 3060 for 1.07x, and my own router managed 0.52x. Plus ten LoRA seeds that all landed at or below the base model, and the llama.cpp pin moving to v0.4.0.
LoRA Skill Compilation Is a Double-Headed Coin Flip
Ten LoRA seeds on identical data spread 3.62 points, against 4.65 for prompt compilation. Not one beat the no-adapter base. Measured on an RTX 3090.
I got the result I wanted. Then I paid $4.97 to run it nine more times.
A frontier model wrote a skill that made our local 27B 10.6 points better. Nine more compilation runs showed the number was noise. Plus the LoRA substrate nobody has measured, MoE routing traced, and 277 GB in a file with no name.
Skills in the Weights: The LoRA Answer to the Prompt Tax
Compiling a skill into a prompt costs 1,383.9 tokens per call, forever. Putting it in a LoRA costs 7.55 GB and a training run. Only one has been measured.
Unsloth Studio Setup Guide: Fine-Tune Qwen 3.5 on Your GPU (Step by Step)
How to install Unsloth Studio, run GGUF models locally, and fine-tune Qwen 3.5 — all in one open-source web UI. Works on Mac, Windows, and Linux.
Fine-Tuning on Mac: LoRA & QLoRA with MLX
Fine-tune Llama, Qwen, and Mistral on Apple Silicon using mlx-lm. Real memory numbers, step-by-step commands, and how to deploy your model with Ollama.
LoRA Training on Consumer Hardware: Fine-Tune Models With 12GB VRAM
QLoRA fine-tunes a 7B model on an RTX 3060 12GB in 2-4 hours. Full Unsloth and Axolotl recipes, VRAM tables, and the GGUF export pipeline.
AI Art Styles & Workflows: SD and Flux Guide
Photorealism, anime, oil painting, concept art, and pixel art on 8GB+ VRAM. Model picks, LoRA stacking at 0.5-0.8 weight, and ComfyUI workflows for each style.
Fine-Tuning LLMs on Consumer Hardware: LoRA and QLoRA Guide
Fine-tune a 7B model on 6-10GB VRAM with QLoRA and Unsloth (2-5x faster, 70% less memory). Only 200-500 examples needed — and in 2026 you can fine-tune a 30B MoE on a single 24GB card. Dataset prep through training on RTX 3060-5090.