Expert Offload
Qwen 3.6 MoE Routing, Measured: Flat Is the Wrong Number
I traced every expert routing decision Qwen 3.6-35B-A3B makes across six workloads on an RTX 3060. Routing isn't flat, and 112 slots is the whole answer.
MoE Offload on RTX 3090: The Curve Is Linear, Not a Cliff (2026)
Every offloaded layer on a 3090 costs about half a millisecond, with no knee anywhere in the curve. Firsthand sweep, the two rules it broke, and a new 26B MoE.
Every SSD-Streaming MoE Engine: What's Real, What's Dead
Eight engines that stream MoE experts from disk appeared in four months. Three have no license file at all. Here's the verified state of each.
Gemma 4 26B in 2GB RAM: The MoE Memory Ladder Explained
One model, three places its experts can live: VRAM, RAM, SSD. We measured the first two on Gemma 4. TurboFieldfare just added the third.
MoE Models Explained: Why Mixtral Uses 46B Parameters But Runs Like 13B
MoE explained with our own 3090 and 3060 numbers: a 35B MoE fits 24 GB and beats dense by up to 4x, offload moved the wall to RAM, and where dense still wins.