Qwen 3.8 35B-A3B Is a Distill: Tested Against Qwen 3.6
📚 More on this topic: Best way to run Qwen 3.6 35B MoE locally · Qwen 3.8 27B vs 3.6 on RTX 3090 · Why Qwen 3.8 27B Feels Slow · Distilled vs Frontier Models · InsiderLLM Benchmarks
You searched for Qwen 3.8 35B-A3B, found a GGUF with a big download count, and you’re about to give it 21 GB of disk. Before you do, here’s the part the repo name skips: Alibaba hasn’t released a Qwen 3.8 35B-A3B. The Qwen 3.8 open line so far is a dense 27B and a 2.4-trillion-parameter flagship nobody runs at home. The 35B you found is empero-ai/Qwen3.8-35B-A3B-Distill. A small outfit called Empero released it on September 16. They took Qwen 3.6 35B-A3B and fine-tuned it on reasoning traces from Qwen 3.8’s big models. So the architecture, the size and the tokenizer are Qwen 3.6’s, and the training on top is Qwen 3.8-flavoured.
That’s a legitimate thing to build, and the licence is Apache 2.0. But it means the right comparison isn’t Qwen 3.8 against Qwen 3.6. It’s this fine-tune against the exact model it started from. If the distill doesn’t beat its own base, the “3.8” in the name is doing marketing work.
So I ran that comparison on my RTX 3090: both models at Q4_K_M, the same llama.cpp build, the same afternoon. The test was my 47-item routing split. Before I downloaded anything I wrote down what would count as a difference.
The short version
| Qwen 3.6 35B-A3B | Qwen 3.8 35B-A3B Distill | |
|---|---|---|
| Routing split, exact (thinking off) | 23 / 47 | 22 / 47 |
| Same, counting only answers that parsed | 23 / 47 | 20 / 47 |
| Messages where it chatted instead of classifying | 0 | 6 |
| Decode, tok/s (empty context / 8K) | 183.9 / 177.6 | 183.6 / 177.4 |
| Routing split, exact (thinking on) | 23 / 47 | 22 / 47 |
| Median reasoning tokens per message (thinking on) | 483 | 94 |
| Messages that hit the 8,192-token cap | 4 | 1 |
| Time for all 47, thinking on | 346 s | 92 s |
| Peak VRAM | 20,980 MiB | 20,980 MiB |
It isn’t smarter on this task. It is just as fast, and it thinks about a fifth as long. And it has a habit the original doesn’t.
What the distill actually is
I read the model card in full before running anything. The parts worth knowing:
- Base:
Qwen/Qwen3.6-35B-A3B, fine-tuned. Same 35B total, 3B active per token. - Method: “SFT (off-policy distillation) on curated teacher traces”. Empero trained it on answers written by bigger models.
- Teachers: “Qwen3.8 2.4T A95B and Qwen3.8 Flash Next”. Those models carry Alibaba’s own licences: qwen3.8-max for the 2.4T and qwen-community-1.0 for Flash Next. Both allow derivatives, with conditions on large commercial and model-as-a-service use. Empero’s weights are Apache 2.0, inherited from the base. If you’re shipping a product on this, read the teacher licences yourself. The card doesn’t mention them.
- The card’s own benchmarks are MMLU and ARC scored by loglikelihood: base 0.838 vs distill 0.834 on MMLU, gains of 3 to 5 points on ARC. Loglikelihood scoring never makes the model write anything. So none of those numbers exercise the thing the distill was trained to change, which is how it reasons when it talks.
- The card’s own warning: the student “produces noticeably shorter outputs than the base”, and long chains of thought “are more likely to be cut short.”
That last line turned out to be the most accurate claim on the card.
Same score: 22 vs 23, and the gate says no difference
The test is the intent classifier from my own agent. Each message gets three labels (which tool, which mode, how far back to look), and an answer counts only if all three match. These are the same 47 items and the same scoring I used for Bonsai 2. Thinking off, temperature 0, 100 output tokens, two passes per model. Both models gave identical answers on both passes.
I wrote the rule before the download: within one item is no difference, three or more is a real result, two is inconclusive. The base got 23, the distill 22. No difference.
Per field it’s even closer than that. The distill got two more tool calls right (31 to 29) and two more modes (29 to 27), and one fewer scope (27 to 28). On 21 items both were right, on 23 both were wrong, and they split the other three. For context, Qwen 3.6 27B dense scores 23 on this split and Qwen 3.8 27B scored 19. The 35B MoE is right with them.
The new habit: it sometimes chats instead of classifying
Here’s the part the score hides.
The base returned valid JSON on all 47 messages. The distill did it on 41. On the other 6 it ignored the system prompt and answered the message as if it were the assistant being spoken to, until it ran out of tokens:
That’s a genuinely good question, and I think the honest answer is: I don’t know.
All six have the same shape. The user is talking to the assistant about itself: its memory, whether it’s the same “you” from one session to the next. The base files those under a label every time. The distill takes the bait and has the conversation.
It doesn’t show up in the headline score because my scorer copies what my production agent does. If the model’s answer won’t parse, it falls back to “just chat” and moves on. On two of the six messages, “just chat” happened to be correct. So the distill’s 22 includes two points it got by failing in a convenient direction. Count only the answers that actually parsed and it’s 20 against 23. That’s three items, which would cross my “real result” line. But I set the gate on the production scoring before the run, so 22 vs 23 is the result and the 20 goes beside it, labelled.
One more thing, because I’d rather say it than have you find it in the record. I wrote a rule in advance for this. Five or more unparseable answers would mean the distill can’t follow thinking-off. The rule fired, but not for the reason I wrote it. I expected leaked reasoning. Nothing leaked; there wasn’t a single think tag in the six. It’s a different failure, and the record says so.
Same speed, to the decimal
These two files are the same architecture and nearly the same size, so I expected identical speed. That’s what I got.
| llama-bench, tok/s | Qwen 3.6 35B-A3B | Distill | difference |
|---|---|---|---|
| decode, empty context | 183.94 | 183.62 | −0.2% |
| decode, 4K deep | 181.08 | 181.03 | −0.0% |
| decode, 8K deep | 177.58 | 177.44 | −0.1% |
| prefill 512, empty context | 3,661 | 3,657 | −0.1% |
| prefill 512, 8K deep | 3,366 | 3,342 | −0.7% |
That’s the mean of three alternating reps, each after a discarded warm-up run. The spread between reps was under 0.6 tok/s on decode. Both peaked at 20,980 MiB on a 24 GB card, so there’s room for context either way.
One thing in the GGUF you might wonder about: the distill ships an extra layer for multi-token prediction, about 0.5 GB of it. Mainline llama.cpp doesn’t load it unless you turn on speculative decoding. It logs those tensors as unused and skips them, which is why VRAM matches to the MiB. I’d expected it to cost VRAM and said so in my plan. It didn’t.
If you have the Qwen 3.6 35B setup running already, swapping in the distill changes nothing about your flags, your fit or your tok/s.
A fifth of the reasoning
This is where the two models really differ. I ran all 47 again with thinking on, up to 8,192 tokens, one pass each.
The scores barely moved: 23 for the base, 22 for the distill, no gate on this one. But the base’s median reasoning was 483 tokens a message. The distill’s was 94. Over the whole set the base spent 56,862 reasoning tokens and the distill 13,728. That works out to 346 seconds against 92 for the same 47 answers at the same decode speed.
The cap hits are telling too. The base ran into the 8,192 limit on four messages. Each time it was deliberating, circling back to a choice it had already made. One ends mid-JSON with the cap. If you’ve read why Qwen 3.8 27B feels slow, it’s the same family trait. The distill hit the cap once, and that was a degenerate loop: the same short line repeated until the limit. The card warns about greedy decoding causing loops, and I ran temperature 0 on purpose so both models got identical settings. At the card’s recommended 0.6 that loop might not happen. I didn’t test it.
So the card’s “shorter outputs” claim holds, and on this task shorter cost nothing in accuracy. Whether it costs anything on long maths or code, the domains the traces were weighted toward, I didn’t measure.
Which one should you run
If you care about latency and short answers, the distill is fine. For chat, quick questions, or anything where you’re waiting on a thinking model to stop thinking, a fifth of the reasoning at the same tok/s is the whole pitch, and it delivered. On a 3090 that averaged about 2 seconds a message against 7.
If something downstream parses the output, stay on Qwen 3.6, or validate. That means agents, tool routers, anything with a JSON contract. Six messages in 47 where the model dropped the format and started talking is a real failure rate for a pipeline. Mine falls back to a safe default. Yours might not. If you do run the distill in an agent, check every response against your schema and retry on failure. Don’t trust it to stay in role on messages that address the assistant directly.
If you want Qwen 3.8, this isn’t it. It’s Qwen 3.6 with some 3.8 habits trained in. The real Qwen 3.8 you can run on one card is the 27B dense, at about a quarter of this model’s decode speed.
Limits
- One task. A 47-item routing split with short outputs. It says nothing about the maths, code and tool-use claims on the card, or about long-form writing, which the card itself flags as a weak spot.
- One quant, one build, one card. Q4_K_M, llama.cpp v0.4.0, RTX 3090.
- The base file’s origin is unrecorded. The Qwen 3.6 Q4_K_M I used was probably quantized locally from my own F16 in July, but no record says so. I hashed it before the run (sha256
203c3a7c…) so the result is tied to that exact file. The distill’s file matched the hash on Hugging Face. Both are plain Q4_K_M with no importance matrix, made by different llama.cpp versions. I used the published files rather than re-quantizing both. That leaves the quantizer as a small uncontrolled variable. - My box was running its memory slow. Tamanna’s DDR4 is at 2133 while I sort out a memory-training problem. Both models sat entirely in VRAM, so system RAM speed doesn’t enter into these numbers.
- Thinking on was one pass at temperature 0, not the card’s 0.6.
Everything is in the record: tamanna-qwen38-distill-vs-qwen36-35b-2026-10-07. It has the pre-registration, every raw completion, the per-item table, hashes, power and clock telemetry, and the memory watchdog log. The six speed rows are in the benchmarks dataset.
A weekly email with every new guide and measured benchmark.
Subscribe — free, no spam