RTX 3060
A 27B for your 12 GB card. It lost one field.
Bonsai 2 27B fits a 12 GB card and decodes 1.8x faster than its Q4, then drops from 19 to 12 of 47 on a routing split, all of it on one field. Plus 23 used-GPU price corrections, the 3060 fair-price ladder, and eight Quick Hits.
Stop writing the answer. Pick it. 47 items waiting.
Codacus's parallel-decision llama.cpp branch answers a schema in one forward pass, and the ten-seeds 47-item split is the ready-made test. Plus ChatGPT-User fetches sliding 5 percent a week while OAI-SearchBot tripled, the v0.4.0 pin on the 3060, and a third bench rig.
Nineteen gigabytes of my 3090 sat empty. The 3060 kept up.
Stock llama.cpp put 5.4 GB of a 177B MoE on an RTX 3090 and a 3060 kept pace; -ncmoe 29 buys 34 to 40 percent and a real 32 GB box does 8.7 tok/s. Plus a 4 GB GTX 1650 at 20 tok/s on a 35B MoE, and the MoE primer corrected in public.
GTX 1650 vs RTX 3060 on a 35B MoE: What the Card Buys
A $60-class GTX 1650 4 GB runs Qwen3.6-35B-A3B at 20 tok/s on 32 GB of RAM. The RTX 3060 in the same slot does 28 at the same setting and 39 tuned. Measured.
Nvidia's router dealt the cards evenly. That was the whole problem.
Nvidia PAIR split 20 requests 10/10 across an RTX 3090 and a 3060 for 1.07x, and my own router managed 0.52x. Plus ten LoRA seeds that all landed at or below the base model, and the llama.cpp pin moving to v0.4.0.
A 177B Model on a 3060: The 32 GB Number Nobody Measured
Qwen3.8-Flash-Next on an RTX 3090 and a 3060. Stock, the 3090 matches the video's 3060. One flag buys 34 to 40 percent. A real 32 GB box: 8.7 tok/s, not 22.
Nvidia PAIR Bought Me 7 Percent. My Own Router Cost Me Half.
Nvidia's new home-network AI router split 20 requests across an RTX 3090 and a 3060 for 1.07x over the 3090, then dropped half a 27B queue. Mine: 0.52x.
I got the result I wanted. Then I paid $4.97 to run it nine more times.
A frontier model wrote a skill that made our local 27B 10.6 points better. Nine more compilation runs showed the number was noise. Plus the LoRA substrate nobody has measured, MoE routing traced, and 277 GB in a file with no name.
Qwen 3.6 MoE Routing, Measured: Flat Is the Wrong Number
I traced every expert routing decision Qwen 3.6-35B-A3B makes across six workloads on an RTX 3060. Routing isn't flat, and 112 slots is the whole answer.
MoE Offload on RTX 3090: The Curve Is Linear, Not a Cliff (2026)
Every offloaded layer on a 3090 costs about half a millisecond, with no knee anywhere in the curve. Firsthand sweep, the two rules it broke, and a new 26B MoE.
Gemma 4 26B in 2GB RAM: The MoE Memory Ladder Explained
One model, three places its experts can live: VRAM, RAM, SSD. We measured the first two on Gemma 4. TurboFieldfare just added the third.
Qwen 35B-A3B on RTX 3090: 157 tok/s With No Offload (2026)
The whole 35B sits on a 24GB card with 2.4 GiB spare, no expert offload. Firsthand numbers, the harness caveat, and why max offload loses to a 3060.
RTX 3060 vs 3060 Ti vs 3070 for Local AI
The RTX 3060 12GB now costs $50 MORE than the 3060 Ti and still wins for LLMs. Why the used market flipped, what $/GB reveals, and when the 3070 makes sense.
Best Used GPUs for Local AI: 2026 Buying Guide
RTX 3090 at ~$1,200-1,400 for 24GB, RTX 3060 12GB at $200-400, RTX 3080 at $350-400. Tier rankings, fair prices, what to avoid (skip the 8GB 3070), and where to buy safely.
Best GPU Under $500 for Local AI (2026 Picks)
Find the best GPU under $500 for running local AI in 2026. RTX 4060 Ti 16GB, used RTX 3080, RTX 3060 12GB, and RX 7700 XT compared with real benchmarks.
Best GPU Under $300 for Local AI (2026 Picks)
Find the best GPU under $300 for local AI. We compare the RTX 3060 12GB, RX 7600, and Intel Arc B580 with VRAM analysis, LLM benchmarks, and real pricing.
Used GPU Buying Guide for Local AI: How to Buy Smart
Used RTX 3060 12GB at $200-400, RTX 3090 24GB at $1,000-1,400 as the 2026 memory shortage bites. Fair ranges, scam red flags, the Ti-badge trap, where to buy safely.
What Can You Actually Run on 12GB VRAM?
Qwen 3.5 9B at Q8_0 runs near-lossless on 12GB, Qwen 2.5 14B at Q4 hits 30 tok/s, and SDXL generates without workarounds. Every model that fits on an RTX 3060 12GB and the best upgrade path.
Used Optiplex + RTX 3060 = Local AI for $500 (Full Build)
Build a local AI PC for about $500: used Dell Optiplex + RTX 3060 12GB runs Qwen 3.5 9B at 35-45 tok/s. Full parts list, where to buy, 2026 pricing.
A weekly email with every new guide and measured benchmark.
Subscribe — free, no spam