Small dense 1B-4B models run at 18-55 tok/s. Qwen3 4B at Q4 is the 4GB sweet spot for chat and simple coding. 7B models don't fit — and the MoE offload trick starts at 12GB, not here.