RTX 3090
A 27B for your 12 GB card. It lost one field.
Bonsai 2 27B fits a 12 GB card and decodes 1.8x faster than its Q4, then drops from 19 to 12 of 47 on a routing split, all of it on one field. Plus 23 used-GPU price corrections, the 3060 fair-price ladder, and eight Quick Hits.
Bonsai 2 vs Qwen3.8-27B on an RTX 3090: Speed Doubled, Scope Broke
Bonsai 2 27B, PrismML's ternary Qwen3.8, on a 3090: 1.8x the decode of Q4 from a 7.2 GB file, then 12 of 47 against 19 on a routing split. One field broke.
What Is Jev, and Can You Run One on Your Own GPU?
Jev picks from options you supply instead of writing, in one pass, with a probability per answer. What it is, what it costs, and the open version for a 3090.
Stop writing the answer. Pick it. 47 items waiting.
Codacus's parallel-decision llama.cpp branch answers a schema in one forward pass, and the ten-seeds 47-item split is the ready-made test. Plus ChatGPT-User fetches sliding 5 percent a week while OAI-SearchBot tripled, the v0.4.0 pin on the 3060, and a third bench rig.
Jev Mode on a 3090: 26 ms per Token You Don't Write
Jev-style parallel decisions in llama.cpp on Qwen3.6-27B, RTX 3090: 1.3x faster at four tokens, 5.5x at fifty, 21 vs 23 of 47. The gain is the tokens you skip.
Nineteen gigabytes of my 3090 sat empty. The 3060 kept up.
Stock llama.cpp put 5.4 GB of a 177B MoE on an RTX 3090 and a 3060 kept pace; -ncmoe 29 buys 34 to 40 percent and a real 32 GB box does 8.7 tok/s. Plus a 4 GB GTX 1650 at 20 tok/s on a 35B MoE, and the MoE primer corrected in public.
Nvidia's router dealt the cards evenly. That was the whole problem.
Nvidia PAIR split 20 requests 10/10 across an RTX 3090 and a 3060 for 1.07x, and my own router managed 0.52x. Plus ten LoRA seeds that all landed at or below the base model, and the llama.cpp pin moving to v0.4.0.
A 177B Model on a 3060: The 32 GB Number Nobody Measured
Qwen3.8-Flash-Next on an RTX 3090 and a 3060. Stock, the 3090 matches the video's 3060. One flag buys 34 to 40 percent. A real 32 GB box: 8.7 tok/s, not 22.
Nvidia PAIR Bought Me 7 Percent. My Own Router Cost Me Half.
Nvidia's new home-network AI router split 20 requests across an RTX 3090 and a 3060 for 1.07x over the 3090, then dropped half a 27B queue. Mine: 0.52x.
LoRA Skill Compilation Is a Double-Headed Coin Flip
Ten LoRA seeds on identical data spread 3.62 points, against 4.65 for prompt compilation. Not one beat the no-adapter base. Measured on an RTX 3090.
I got the result I wanted. Then I paid $4.97 to run it nine more times.
A frontier model wrote a skill that made our local 27B 10.6 points better. Nine more compilation runs showed the number was noise. Plus the LoRA substrate nobody has measured, MoE routing traced, and 277 GB in a file with no name.
Skills in the Weights: The LoRA Answer to the Prompt Tax
Compiling a skill into a prompt costs 1,383.9 tokens per call, forever. Putting it in a LoRA costs 7.55 GB and a training run. Only one has been measured.
I Got the AI Result I Wanted. Then I Ran It Nine More Times
A frontier model read my logs and wrote a skill that beat my local 27B baseline by 10.6 points. Nine more compilation runs showed the number was fake.
Ornith 1.5 35B vs Qwen 3.6 on RTX 3090: Speed Tested
Firsthand A-B-B-A bench of Ornith 1.5-35B-A3B against Qwen 3.6-35B-A3B on one RTX 3090. Generation, prefill, VRAM, and the noise floor under all three.
Qwen 3.8 isn't slow. It's just very, very thorough.
Qwen 3.8-27B spent 14,953 tokens on a line the same file writes in nine. All 164 HumanEval problems measured: 92.8% of output is thinking. Plus the four runs that tie it with 3.6.
Why Qwen 3.8 27B Feels Slow: Reasoning Tokens Measured
Qwen 3.8 27B generates at full speed on a 3090 and still crawls. Four runs, two models, two seeds: 92.8% of output is thinking, and the spread runs 320x to 542x.
Qwen 3.8 27B vs 3.6 on RTX 3090: Speed and Quality Tested
Firsthand benchmarks of Qwen 3.8-27B against 3.6-27B on one RTX 3090: generation within a percent, VRAM +254 MiB, and HumanEval pass@1 a statistical tie.
MoE Offload on RTX 3090: The Curve Is Linear, Not a Cliff (2026)
Every offloaded layer on a 3090 costs about half a millisecond, with no knee anywhere in the curve. Firsthand sweep, the two rules it broke, and a new 26B MoE.
Qwen 35B-A3B on RTX 3090: 157 tok/s With No Offload (2026)
The whole 35B sits on a 24GB card with 2.4 GiB spare, no expert offload. Firsthand numbers, the harness caveat, and why max offload loses to a 3060.
Kimi K3 & Qwen 3.8: Open Weights You Can't Run (2026)
Kimi K3's 2.8T weights need 64 accelerators to load. Qwen 3.8's 27B did ship, Apache 2.0, and fits 24GB. Openness and runnability are separate axes.
Inkling 975B vs Your 3090: The Real Memory Math (2026)
Inkling's smallest 1-bit quant is 270GB. A maxed consumer desktop holds 256GB. The open frontier left consumer hardware behind. Here's the honest math.
Qwen 3.7's open weights are overdue — by the math, not vibes
Qwen's own release cadence says the 3.7 open weights should already be out, and they're not. Plus GLM-5.2 running locally: a frontier open model that takes serious hardware.
Best 24GB Backend Shootout: ik_llama vs BeeLlama vs llama.cpp
ik_llama and BeeLlama both finish in 22-23s on the am17an 9-prompt harness vs mainline llama.cpp's 37s — 1.66x and 1.62x speedups via opposite strategies.
Wicked Fast Qwen 3.6 27B: 60 tok/s with MTP on RTX 3090 (2026)
Firsthand bench: 60 tok/s on Qwen 3.6 27B Q4_K_M with MTP on a single RTX 3090 — 1.86x wall-clock speedup over baseline. PR #22673 progress May 6 → May 19.
Wicked Fast Gemma 4 vs Qwen 3.6 on RTX 3090: 3.10x Tested
Same RTX 3090, same llama.cpp build, same bench. Gemma 4 26B-A4B Q4_K_XL: 128 tok/s mean. Qwen 3.6-27B Q4_K_M: 41 tok/s. 3.10x faster, firsthand.
DFlash vs MTP on RTX 3090: I Tested Both Locally
Firsthand head-to-head bench of DFlash + DDTree against MTP (PR #22673) on a single RTX 3090, same Qwen 3.6-27B target. Real numbers, both backends.
How to Fix Slow Qwen 3.6 27B on RTX 3090 (10-80 tok/s)
Qwen 3.6-27B at 12 tok/s on a 3090 when others report 35? The 8-step diagnostic checklist for offload, quants, templates, power limits, and backend choice.
How to Get 2.5x Faster Qwen on RTX 3090 (Free)
I built DFlash on my RTX 3090 and ran the full bench. Real 2.5x speedup on Qwen 3.5 and 3.6 — below the 3.43x README claim, still huge. Here's how.
Best Way to Run Qwen 3.6 35B MoE Locally: VRAM, Speed, Setup
Qwen 3.6-35B-A3B has 35B total params but only 3B active per token. Real tok/s on RTX 3090, 4090, 5070 Ti, dual 5060 Ti, and M3 Ultra. Quants and setup.
Best Way to Get 2x Token Output on RTX 3090: Qwen 3.6 + DFlash
Luce DFlash + DDTree pushes Qwen 3.6-27B Q4_K_M from 35 tok/s to 69 tok/s on a single RTX 3090. Real benchmarks, setup, and honest limits.
RTX 5090 Benchmarks: 5090 vs 4090 vs Used 3090 (2026)
5090 community benches across 4K-131K context, prompt-processing tables, 5090-vs-4090 upgrade math, and InsiderLLM's firsthand 3090 honest-value anchor.
RTX 4090 vs Used RTX 3090 for Local AI: Which to Buy in 2026
Both have 24GB VRAM. One costs about twice as much. RTX 4090 vs used RTX 3090 — real benchmarks, real prices, and who should buy which for local AI.
Best Dual-GPU Local AI Setup: RTX 3090, 5060 Ti (2026)
Dual RTX 3090, 2x RTX 5060 Ti, 2x 2080 Ti modded, mixed setups: real configs for Qwen 3.6, MoE, 70B. Tensor vs pipeline parallelism, llama.cpp/vLLM.
RTX 3090 vs 4070 Ti Super for Local LLMs
Head-to-head comparison of the RTX 3090 and RTX 4070 Ti Super for running LLMs locally. Covers VRAM, speed, power, price, and which to buy for your use case.
Best Used GPUs for Local AI: 2026 Buying Guide
RTX 3090 at ~$1,200-1,400 for 24GB, RTX 3060 12GB at $200-400, RTX 3080 at $350-400. Tier rankings, fair prices, what to avoid (skip the 8GB 3070), and where to buy safely.
Used GPU Buying Guide for Local AI: How to Buy Smart
Used RTX 3060 12GB at $200-400, RTX 3090 24GB at $1,000-1,400 as the 2026 memory shortage bites. Fair ranges, scam red flags, the Ti-badge trap, where to buy safely.
What Can You Actually Run on 24GB VRAM?
Qwen 3.5 27B at Q4 fits in 17GB with 64K+ context. 70B at Q3 with limited context. Flux at full FP16. RTX 3090 at $1,300 vs 4090 at $2,860—every model that fits and which GPU to buy.
Used RTX 3090 Buying Guide for Local AI
24GB VRAM for ~$1,200-1,400 used (Aug 2026)—still the cheapest 24GB card on the market. eBay red flags, PSU requirements (850W minimum), and how to test before your return window closes.
A weekly email with every new guide and measured benchmark.
Subscribe — free, no spam