GTX 1650
Nineteen gigabytes of my 3090 sat empty. The 3060 kept up.
Stock llama.cpp put 5.4 GB of a 177B MoE on an RTX 3090 and a 3060 kept pace; -ncmoe 29 buys 34 to 40 percent and a real 32 GB box does 8.7 tok/s. Plus a 4 GB GTX 1650 at 20 tok/s on a 35B MoE, and the MoE primer corrected in public.
GTX 1650 vs RTX 3060 on a 35B MoE: What the Card Buys
A $60-class GTX 1650 4 GB runs Qwen3.6-35B-A3B at 20 tok/s on 32 GB of RAM. The RTX 3060 in the same slot does 28 at the same setting and 39 tuned. Measured.
What Can You Actually Run on 4GB VRAM?
Small dense 1B-4B models run at 18-55 tok/s. Qwen3 4B at Q4 is the 4GB sweet spot for chat and simple coding. 7B models don't fit — and the MoE offload trick starts at 12GB, not here.
A weekly email with every new guide and measured benchmark.
Subscribe — free, no spam