Cpu-Inference
Qwen 3.8 isn't slow. It's just very, very thorough.
Qwen 3.8-27B spent 14,953 tokens on a line the same file writes in nine. All 164 HumanEval problems measured: 92.8% of output is thinking. Plus the four runs that tie it with 3.6.
The $36 RAM Fix That Made CPU Inference 56% Faster
Adding a second RAM stick to a mini PC lifted CPU token generation 52-58% across four models. Prompt processing moved under 2%. Measured before and after.