HumanEval
Qwen 3.8 isn't slow. It's just very, very thorough.
Qwen 3.8-27B spent 14,953 tokens on a line the same file writes in nine. All 164 HumanEval problems measured: 92.8% of output is thinking. Plus the four runs that tie it with 3.6.
Qwen 3.8 27B vs 3.6 on RTX 3090: Speed and Quality Tested
Firsthand benchmarks of Qwen 3.8-27B against 3.6-27B on one RTX 3090: generation within a percent, VRAM +254 MiB, and HumanEval pass@1 a statistical tie.