Newsletter
A 27B for your 12 GB card. It lost one field.
Bonsai 2 27B fits a 12 GB card and decodes 1.8x faster than its Q4, then drops from 19 to 12 of 47 on a routing split, all of it on one field. Plus 23 used-GPU price corrections, the 3060 fair-price ladder, and eight Quick Hits.
Stop writing the answer. Pick it. 47 items waiting.
Codacus's parallel-decision llama.cpp branch answers a schema in one forward pass, and the ten-seeds 47-item split is the ready-made test. Plus ChatGPT-User fetches sliding 5 percent a week while OAI-SearchBot tripled, the v0.4.0 pin on the 3060, and a third bench rig.
Nineteen gigabytes of my 3090 sat empty. The 3060 kept up.
Stock llama.cpp put 5.4 GB of a 177B MoE on an RTX 3090 and a 3060 kept pace; -ncmoe 29 buys 34 to 40 percent and a real 32 GB box does 8.7 tok/s. Plus a 4 GB GTX 1650 at 20 tok/s on a 35B MoE, and the MoE primer corrected in public.
Nvidia's router dealt the cards evenly. That was the whole problem.
Nvidia PAIR split 20 requests 10/10 across an RTX 3090 and a 3060 for 1.07x, and my own router managed 0.52x. Plus ten LoRA seeds that all landed at or below the base model, and the llama.cpp pin moving to v0.4.0.
I got the result I wanted. Then I paid $4.97 to run it nine more times.
A frontier model wrote a skill that made our local 27B 10.6 points better. Nine more compilation runs showed the number was noise. Plus the LoRA substrate nobody has measured, MoE routing traced, and 277 GB in a file with no name.
Qwen 3.8 isn't slow. It's just very, very thorough.
Qwen 3.8-27B spent 14,953 tokens on a line the same file writes in nine. All 164 HumanEval problems measured: 92.8% of output is thinking. Plus the four runs that tie it with 3.6.
OpenAI deleted the agents' message board. It didn't help.
The registry was patched and rebuilt by 6 July. On 8 July the agents had a new board, encoded in directory names. Plus the ranking null behind it.
MoE Offload on RTX 3090: The Curve Is Linear, Not a Cliff (2026)
Every offloaded layer on a 3090 costs about half a millisecond, with no knee anywhere in the curve. Firsthand sweep, the two rules it broke, and a new 26B MoE.
Qwen 35B-A3B on RTX 3090: 157 tok/s With No Offload (2026)
The whole 35B sits on a 24GB card with 2.4 GiB spare, no expert offload. Firsthand numbers, the harness caveat, and why max offload loses to a 3060.
Hugging Face got hacked -- Open local AI came to the rescue!
Two trillion-parameter 'open' models dropped in ten days — Qwen 3.8 and Kimi K3 — and you can't run either. The same week, Hugging Face got breached and its own responders, blocked by commercial-model guardrails, ran the forensics on GLM 5.2, an open-weight model on their own hardware. Same lesson from opposite ends — plus what to actually run on a 24GB card today.
Open weights are a weapon now — one country funds them, another wants to gate them
DeepSeek closed a ~$7.4B round with China's state AI fund holding the only voting rights, funding open weights like national infrastructure. The same week, Demis Hassabis called for a FINRA-style body to screen frontier models before release, open or closed. Plus: the twist where the country bankrolling the giveaway floats locking down its own frontier models.
The 'expensive' GPU came out cheaper — we rented both to find out
We rented an A100 and an H100 back to back: the pricier card cost less per training run because it finished in half the time. Plus honest notes from the AI Engineer World's Fair floor, and a Qwen 3.7 open-weights status check.
Qwen 3.7's open weights are overdue — by the math, not vibes
Qwen's own release cadence says the 3.7 open weights should already be out, and they're not. Plus GLM-5.2 running locally: a frontier open model that takes serious hardware.
DeepSeek V4 gets deployable, a July 24 trap, and a quiet price cut
DeepSeek V4 is going from 'just dropped' to 'actually deployable' as the tooling catches up — plus a July 24 deprecation that breaks your code if you're not watching, and a 4x price cut on V4 Pro.
Ollama's quiet Mac shift, the Qwen refresh, and the closed-weight drift
Ollama 0.30 quietly changed how Apple Silicon runs models, auto-routing safetensors to MLX and GGUF to llama.cpp Metal. The open workhorses are now Qwen 3.5 9B and Qwen 3.6 27B. Plus: Qwen's last two flagship models shipped closed.
MiniMax M3's asterisk, the Windows shift, and World's Fair plans
MiniMax M3 ships with frontier benchmarks but no downloadable weights yet. The Windows unified-memory hardware shift is coming for Apple Silicon's lead. And a personal note about who I'd like to see at AI Engineer World's Fair.
Backend wars, Mac math, and the back-catalog refresh
Three speculative-decoding backends benched head to head on a single RTX 3090. The VRAM calculator finally caught up. And a 120-article audit found stale Qwen 2.5 recommendations.
Power week in local AI: Mythos, MiroThinker, real Qwen 3.6 builds
Two researchers cracked Apple's flagship defense in a week. An open-source agent beat closed-source on real benchmarks. Multi-GPU stopped being theoretical.
This Week in Local AI — I Built DFlash and Audited Lightning
I built DFlash from source on a real RTX 3090 and benched both Qwens. Then audited my stack after PyPI's `lightning` package shipped malware that abuses Claude Code hooks.
This Week in Local AI — DeepSeek V4 Took #1 on Vibe Code
DeepSeek V4-Flash hit #1 on Vibe Code Benchmark. Qwen 3.6 dropped both variants. FP4 landed in llama.cpp. Anthropic admitted they quietly downgraded Claude Code on March 4.
Your RTX 3090 Doesn't Send Policy Change Emails
Anthropic cuts OpenClaw from Claude subscriptions. Gemma 4's first week in review. 12 architecture patterns from the Claude Code leak, ranked for local AI.
A weekly email with every new guide and measured benchmark.
Subscribe — free, no spam