Benchmarks
Stop writing the answer. Pick it. 47 items waiting.
Codacus's parallel-decision llama.cpp branch answers a schema in one forward pass, and the ten-seeds 47-item split is the ready-made test. Plus ChatGPT-User fetches sliding 5 percent a week while OAI-SearchBot tripled, the v0.4.0 pin on the 3060, and a third bench rig.
Nineteen gigabytes of my 3090 sat empty. The 3060 kept up.
Stock llama.cpp put 5.4 GB of a 177B MoE on an RTX 3090 and a 3060 kept pace; -ncmoe 29 buys 34 to 40 percent and a real 32 GB box does 8.7 tok/s. Plus a 4 GB GTX 1650 at 20 tok/s on a 35B MoE, and the MoE primer corrected in public.
Nvidia's router dealt the cards evenly. That was the whole problem.
Nvidia PAIR split 20 requests 10/10 across an RTX 3090 and a 3060 for 1.07x, and my own router managed 0.52x. Plus ten LoRA seeds that all landed at or below the base model, and the llama.cpp pin moving to v0.4.0.
I got the result I wanted. Then I paid $4.97 to run it nine more times.
A frontier model wrote a skill that made our local 27B 10.6 points better. Nine more compilation runs showed the number was noise. Plus the LoRA substrate nobody has measured, MoE routing traced, and 277 GB in a file with no name.
Qwen 3.6 MoE Routing, Measured: Flat Is the Wrong Number
I traced every expert routing decision Qwen 3.6-35B-A3B makes across six workloads on an RTX 3060. Routing isn't flat, and 112 slots is the whole answer.
The $36 RAM Fix That Made CPU Inference 56% Faster
Adding a second RAM stick to a mini PC lifted CPU token generation 52-58% across four models. Prompt processing moved under 2%. Measured before and after.
MoE Offload on RTX 3090: The Curve Is Linear, Not a Cliff (2026)
Every offloaded layer on a 3090 costs about half a millisecond, with no knee anywhere in the curve. Firsthand sweep, the two rules it broke, and a new 26B MoE.
Qwen 35B-A3B on RTX 3090: 157 tok/s With No Offload (2026)
The whole 35B sits on a 24GB card with 2.4 GiB spare, no expert offload. Firsthand numbers, the harness caveat, and why max offload loses to a 3060.
How to Get 2.5x Faster Qwen on RTX 3090 (Free)
I built DFlash on my RTX 3090 and ran the full bench. Real 2.5x speedup on Qwen 3.5 and 3.6 — below the 3.43x README claim, still huge. Here's how.
RTX 5090 Benchmarks: 5090 vs 4090 vs Used 3090 (2026)
5090 community benches across 4K-131K context, prompt-processing tables, 5090-vs-4090 upgrade math, and InsiderLLM's firsthand 3090 honest-value anchor.
LM Studio vs llama.cpp: Why Your Model Runs Slower in the GUI
LM Studio uses llama.cpp under the hood but often runs 30-50% slower. Bundled runtime lag, UI overhead, and default settings explain the gap. How to benchmark it yourself and when the convenience is worth it.
RTX 5060 Ti Review for Local AI — The New Budget King
Real benchmarks for the RTX 5060 Ti 16GB running local LLMs. Qwen 3.5 35B at 44 tok/s, 100K context for ~$430. Compared against RTX 3060, 3090, and 4060 Ti.
The Benchmarks Lie: Why LLM Scores Don't Predict Real-World Performance
MMLU scores drop 14-17 points when contamination is removed. HumanEval is saturated at 94%. Models trained on the test set. Here's what to measure instead.
Distilled vs Frontier Models for Local AI — What You're Actually Getting
That local model you love was probably trained on stolen outputs from Claude or GPT. Here's what distillation actually does to a model's reasoning, where it breaks, and why it matters most for agentic work.
Best Qwen 3.5 Setup: When to Stay vs Move to 3.6 (2026)
3.5 is Qwen's stable open workhorse — 3.6 replaced only two tiers, 3.7 went closed. Which 3.5 model on which GPU, when to stay vs move to 3.6.
Best LLM Speed Trick: ExLlamaV2 vs llama.cpp Benchmarks (50-85% Faster)
Head-to-head speed benchmarks on RTX 3090 and 4090. ExLlamaV2 generates tokens 50-85% faster than llama.cpp on NVIDIA GPUs. Full comparison with setup guides for both.
A weekly email with every new guide and measured benchmark.
Subscribe — free, no spam