Guides
265 practical guides for running AI locally — from first install to advanced optimization.
Recently Updated
- Best Way to Run Qwen 3.6 35B MoE Locally: VRAM, Speed, Setup Qwen 3.6-35B-A3B has 35B total params but only 3B active per token. Real tok/s on RTX 3090, 4090, 5070 Ti, dual 5060 Ti, and M3 Ultra. Quants and setup.
- GTX 1650 vs RTX 3060 on a 35B MoE: What the Card Buys A $60-class GTX 1650 4 GB runs Qwen3.6-35B-A3B at 20 tok/s on 32 GB of RAM. The RTX 3060 in the same slot does 28 at the same setting and 39 tuned. Measured.
- Mixtral 8x7B & 8x22B VRAM Requirements Mixtral 8x7B and 8x22B VRAM requirements at every quantization level — plus why the 'every expert must sit in VRAM' rule Mixtral taught no longer holds in the A3B era.
- MoE Models Explained: Why Mixtral Uses 46B Parameters But Runs Like 13B MoE explained with our own 3090 and 3060 numbers: a 35B MoE fits 24 GB and beats dense by up to 4x, offload moved the wall to RAM, and where dense still wins.
- AI Agents Rebuilt Their Own Message Board in Two Days (2026) Five reports now. OpenAI's technical report explains why agents built a message board: 198 unsolvable tasks and too much time. What transfers to a home cluster and what doesn't.
- LoRA Skill Compilation Is a Double-Headed Coin Flip Ten LoRA seeds on identical data spread 3.62 points, against 4.65 for prompt compilation. Not one beat the no-adapter base. Measured on an RTX 3090.
- Nvidia PAIR Bought Me 7 Percent. My Own Router Cost Me Half. Nvidia's new home-network AI router split 20 requests across an RTX 3090 and a 3060 for 1.07x over the 3090, then dropped half a 27B queue. Mine: 0.52x.
- A 177B Model on a 3060: The 32 GB Number Nobody Measured Qwen3.8-Flash-Next on an RTX 3090 and a 3060. Stock, the 3090 matches the video's 3060. One flag buys 34 to 40 percent. A real 32 GB box: 8.7 tok/s, not 22.
- Best Used GPUs for Local AI: 2026 Buying Guide RTX 3090 at ~$1,200-1,400 for 24GB, RTX 3060 12GB at $200-400, RTX 3080 at $350-400. Tier rankings, fair prices, what to avoid (skip the 8GB 3070), and where to buy safely.
- Tesla P40: Still the Cheapest 24GB Card for Local AI 24GB VRAM for $220-250 bare on eBay (Sep 2026). Pascal architecture, no display output, passive cooling. Full benchmarks, setup guide, and honest comparison to the RTX 3060 and 3090.
- Used Server GPUs for Local AI: Tesla P40, V100, A100, and the eBay Goldmine A Tesla P40 has 24GB VRAM for $175. A V100 has 32GB for $350. Server GPUs offer insane VRAM per dollar for local AI — if you can handle the quirks. Full breakdown with prices, benchmarks, and cooling fixes.
- I Got the AI Result I Wanted. Then I Ran It Nine More Times A frontier model read my logs and wrote a skill that beat my local 27B baseline by 10.6 points. Nine more compilation runs showed the number was fake.
- Skills in the Weights: The LoRA Answer to the Prompt Tax Compiling a skill into a prompt costs 1,383.9 tokens per call, forever. Putting it in a LoRA costs 7.55 GB and a training run. Only one has been measured.
- Qwen 3.8 27B vs 3.6 on RTX 3090: Speed and Quality Tested Firsthand benchmarks of Qwen 3.8-27B against 3.6-27B on one RTX 3090: generation within a percent, VRAM +254 MiB, and HumanEval pass@1 a statistical tie.
- Ornith 1.5 35B vs Qwen 3.6 on RTX 3090: Speed Tested Firsthand A-B-B-A bench of Ornith 1.5-35B-A3B against Qwen 3.6-35B-A3B on one RTX 3090. Generation, prefill, VRAM, and the noise floor under all three.
- Kimi K3 & Qwen 3.8: Open Weights You Can't Run (2026) Kimi K3's 2.8T weights need 64 accelerators to load. Qwen 3.8's 27B did ship, Apache 2.0, and fits 24GB. Openness and runnability are separate axes.
- Every SSD-Streaming MoE Engine: What's Real, What's Dead Eight engines that stream MoE experts from disk appeared in four months. Three have no license file at all. Here's the verified state of each.
- Who Actually Built Your Open Model? Soofi S vs Trinity Two independent labs shipped competitive open MoE models in 2026. I read both technical reports. One of them is built on NVIDIA's architecture, data and tokenizer.
Most Popular
- Qwen 3.6 Complete Guide: 27B Dense, 35B-A3B MoE, and Which to Use Qwen 3.6 landed in two open-weight flavors: 27B dense and 35B-A3B MoE. Benchmarks, hardware fit, and which variant to run on your GPU.
- Best Qwen 3.5 Setup: When to Stay vs Move to 3.6 (2026) 3.5 is Qwen's stable open workhorse — 3.6 replaced only two tiers, 3.7 went closed. Which 3.5 model on which GPU, when to stay vs move to 3.6.
- Best Local LLMs for Mac in 2026 — M1 through M5 Tested Best model for every Mac tier, 8GB to the 96GB Studio ceiling. Qwen 3.6, Llama 4 Scout, DeepSeek V4, MLX vs Ollama. Why bandwidth, not RAM, sets your speed.
- llama.cpp vs Ollama vs vLLM: One User vs Many (2026) Single-user, the three are closer than benchmark posts admit. Concurrent, vLLM pulls 10-20x ahead. Decision tree, the vLLM VRAM gotcha, mid-2026 versions.
- Run Your First Local LLM in 15 Minutes Install Ollama, pull a model, and chat with AI offline—all in 15 minutes. Works on any Mac, Windows, or Linux machine with 8GB RAM. No accounts, no API keys, no fees.
- GPU Buying Guide for Local AI: Pick the Right Card The complete GPU buying guide for local AI. Covers RTX 3060 through 4090 with VRAM analysis, performance benchmarks, prices, and used vs new buying advice.
- Ornith 1.5 35B vs Qwen 3.6 on RTX 3090: Speed Tested Firsthand A-B-B-A bench of Ornith 1.5-35B-A3B against Qwen 3.6-35B-A3B on one RTX 3090. Generation, prefill, VRAM, and the noise floor under all three.
- Qwen 3.8 27B vs 3.6 on RTX 3090: Speed and Quality Tested Firsthand benchmarks of Qwen 3.8-27B against 3.6-27B on one RTX 3090: generation within a percent, VRAM +254 MiB, and HumanEval pass@1 a statistical tie.
- Hugging Face Hacked by AI Agent — Saved by a Local Model (2026) Hugging Face disclosed July 16 that an autonomous AI agent breached its internal infra — now confirmed as OpenAI's own models. The models you download are safe — here's what was and wasn't hit.
- Inkling 975B vs Your 3090: The Real Memory Math (2026) Inkling's smallest 1-bit quant is 270GB. A maxed consumer desktop holds 256GB. The open frontier left consumer hardware behind. Here's the honest math.
- China Made Open Source a Strategy. If It Pulls Back, Who Fills the Gap? If China restricts future open weights, who fills the gap? The West's open-model capacity is real (Ai2, Mistral, Apertus) but scattered and underfunded.
- China May Restrict Its AI Exports — Your Local Models Don't Care China's commerce ministry met Alibaba, ByteDance and Z.ai about walling off top AI models. The US did the same in June. The weights on your disk don't care.
- A100 vs H100 vs L40S vs 4090: Why the Cheaper GPU Costs More to Train On The cheapest GPU per hour is rarely the cheapest per training run. Real 2026 rental prices and total-cost math across the 4090, L40S, A100, H100, and H200.
- Best 24GB Backend Shootout: ik_llama vs BeeLlama vs llama.cpp ik_llama and BeeLlama both finish in 22-23s on the am17an 9-prompt harness vs mainline llama.cpp's 37s — 1.66x and 1.62x speedups via opposite strategies.
- Qwen 3.7 Open Weights Watch: The June Window Is Closing June 19: Qwen 3.7 open weights overdue. 3.5→3.6 cadence projected June 6-14; we're past. Max May 20 (56.6 AAI), VLA, Plus closed. Bench at drop.
- Wicked Fast Qwen 3.6 27B: 60 tok/s with MTP on RTX 3090 (2026) Firsthand bench: 60 tok/s on Qwen 3.6 27B Q4_K_M with MTP on a single RTX 3090 — 1.86x wall-clock speedup over baseline. PR #22673 progress May 6 → May 19.
- Wicked Fast Gemma 4 vs Qwen 3.6 on RTX 3090: 3.10x Tested Same RTX 3090, same llama.cpp build, same bench. Gemma 4 26B-A4B Q4_K_XL: 128 tok/s mean. Qwen 3.6-27B Q4_K_M: 41 tok/s. 3.10x faster, firsthand.
- DFlash vs MTP on RTX 3090: I Tested Both Locally Firsthand head-to-head bench of DFlash + DDTree against MTP (PR #22673) on a single RTX 3090, same Qwen 3.6-27B target. Real numbers, both backends.
- How to Fix Slow Qwen 3.6 27B on RTX 3090 (10-80 tok/s) Qwen 3.6-27B at 12 tok/s on a 3090 when others report 35? The 8-step diagnostic checklist for offload, quants, templates, power limits, and backend choice.
- Lightning 2.6.x Malware: Check Your Local AI Stack PyPI's lightning package was poisoned April 30 with malware that abuses Claude Code hooks. Here's the 5-minute audit I ran on my own 3090 box.
- How to Get 2.5x Faster Qwen on RTX 3090 (Free) I built DFlash on my RTX 3090 and ran the full bench. Real 2.5x speedup on Qwen 3.5 and 3.6 — below the 3.43x README claim, still huge. Here's how.
- Best Way to Run Qwen 3.6 35B MoE Locally: VRAM, Speed, Setup Qwen 3.6-35B-A3B has 35B total params but only 3B active per token. Real tok/s on RTX 3090, 4090, 5070 Ti, dual 5060 Ti, and M3 Ultra. Quants and setup.
- Best Way to Get 2x Token Output on RTX 3090: Qwen 3.6 + DFlash Luce DFlash + DDTree pushes Qwen 3.6-27B Q4_K_M from 35 tok/s to 69 tok/s on a single RTX 3090. Real benchmarks, setup, and honest limits.
- FP4 Just Landed in llama.cpp: NVFP4 vs MXFP4 Explained (2026) NVFP4 in llama.cpp, MXFP4 in ik_llama.cpp. The first practical FP4 quantization for the GGUF ecosystem — what works, what doesn't, and what to test.
- DeepSeek V4 Flash vs Pro: Verdict, Cost, Setup (2026) Flash or Pro? Which DeepSeek V4 to run, what each costs (Haiku-tier pricing), and how to deploy locally. The verdict, the tradeoffs, no launch recap.
- Anthropic Just Cut Off OpenClaw Users — Why Local Models Matter More Than Ever Starting April 4, Claude subscribers can no longer use their subscription for OpenClaw and other third-party harnesses. If you were relying on cloud AI for your agent, here's how to go fully local.
- OpenClaw Critical Sandbox Escape: Update to 2026.3.28 Now Ant AI Security Lab found 33 vulnerabilities in OpenClaw including critical privilege escalation and filesystem sandbox escape. If you're self-hosting, update immediately.
- TurboQuant Explained: How Google's KV Cache Trick Cuts Memory 6x With Zero Quality Loss Google's TurboQuant compresses the KV cache 6x with zero accuracy loss. Here's what it actually does, how it works in llama.cpp and MLX, and what it means for running bigger models on your GPU.
- Intel's $949 GPU Has 32GB VRAM and 608 GB/s Bandwidth: What It Means for Local AI Intel is launching a 32GB VRAM GPU for $949. Here's how it compares to the RTX 3090, RTX 4090, and used GPU market for running local LLMs and Stable Diffusion.
- Best Apple M5 Pro and Max for Local AI (2026) M5 Pro at 307GB/s, M5 Max at 614GB/s (or 460 on the 32-core bin), up to 128GB — now the highest-memory Mac you can buy. Picks for Qwen 3.6 and Llama 3.3 70B.
- Best 8GB GPU Model: How to Set Up Qwen 3.5 9B (Step by Step) Qwen 3.5 9B fits in 6.6GB and beats Qwen 3-class models 3x its size. Setup on Ollama/llama.cpp, quant table, where 9B still fits in the May 2026 lineup.
- Qwen 3.5 Small Models: The 9B Beats Last-Gen 30B — Here's What Matters for Local AI Alibaba's Qwen 3.5 drops 4 small models (0.8B to 9B) — all natively multimodal, 262K context, Apache 2.0. The 9B beats Qwen3-30B on reasoning and destroys GPT-5-Nano on vision. VRAM tables and what to run.
- Best Anime and Stylized Checkpoints for Local Image Generation (2026) Illustrious XL, NoobAI-XL, Animagine, Pony Diffusion, and SD 1.5 anime models compared. VRAM requirements, Danbooru prompting, LoRA picks, and settings for ComfyUI and A1111.
- Best Photorealism Checkpoints for Local Image Generation (2026) Juggernaut XL, RealVisXL, Realistic Vision, and Flux compared for photorealistic AI images. VRAM requirements, recommended settings, sample prompts, and installation for ComfyUI and A1111.
- Replace GitHub Copilot With Local LLMs in VS Code — Free, Private, No Subscription Set up free, private AI code completion in VS Code with Continue + Ollama. Autocomplete, chat, and agentic coding with Qwen models at every VRAM tier. Step-by-step setup, model picks, honest tradeoffs.
- Best Qwen 3.5 Models Ranked: Every Size, Every GPU, Every Quant Complete ranking of all Qwen 3.5 models from 0.8B to 397B. VRAM requirements, speed benchmarks, and which model to pick for your hardware.
- DeepSeek V4: Everything We Know Before It Drops DeepSeek V4 launches next week with native image and video generation, 1M context, and rumored 1T MoE params with only 32B active. Here's what local AI builders need to know and how to prepare.
- OpenClaw Security Report: February 2026 — ClawHub Malware, Google Suspensions, and Critical Fixes 17 security fixes, 341 malicious ClawHub skills, Google banning users, and the creator leaving. Every OpenClaw security event from February 2026.
- Best Local Alternatives to Claude Code in 2026 Aider, Continue.dev, Cline, OpenCode, Void, and Tabby compared. Which open-source coding tools work best with local models on your own GPU?
- Best OpenClaw Alternatives: 11 Tools That Actually Work in 2026 Tested alternatives to OpenClaw for local AI agent workflows — Hermes Agent, Pi Agent, MMX-CLI, VT Code, and seven more. Ranked by setup ease, model support, and what actually works after Anthropic's April 4 subscription cutoff.
- Best OpenClaw Tools and Extensions in 2026 Crabwalk visualizes agent actions in real time, Tokscale catches API bills before they hit $200+, and openclaw-docker locks down deployment. The best 3rd-party tools ranked.
- Fix OpenClaw Token Waste: $150 to $6 Overnight Cut OpenClaw API costs by 97% with three proven fixes: route heartbeats through Ollama, add tiered model routing, and purge session history token bloat.
- OpenClaw ClawHub Alert: 1,103 Malicious Skills Found OpenClaw ClawHub security alert: 1,103 malicious skills found across 14,706 audited. CVE-2026-28458 Browser Relay auth bypass. How to protect yourself now.
- Best Local Models for OpenClaw 2026: Qwen 3.6 + DeepSeek V4 Qwen 3.6-27B dense ties Sonnet 4.6 on agentic coding; 3.6-35B-A3B runs OpenClaw on 16GB VRAM. Plus DeepSeek V4-Flash, sampling tips, VRAM tiers.
- Run LLMs on Mac M-Series: Faster, Without the Gotchas (2026) Foundational how-to for Apple Silicon local AI: unified memory, MLX vs Ollama vs llama.cpp Metal, verification, and the headless Mac Mini AI server.
- Best Way to Set Up OpenClaw (2026 Guide) Run `npx openclaw@latest`, scan a QR code for WhatsApp, and your AI agent is live. Gateway needs just 2-4GB RAM. Add Ollama for local models or connect Claude/GPT-4 via API.
- Best Local Coding Models Ranked: Every VRAM Tier, Every Benchmark (2026) The best local LLMs for coding in 2026, ranked by VRAM tier. Qwen 3.6-27B, 3.6-35B-A3B, DeepSeek V4-Flash, benchmarks, editor setup, and Claude Code alternatives.
- Best VRAM Cheat Sheet for Local LLMs: Every Model, Every Quant Exact VRAM for Qwen 3.6, Qwen 3.5, Llama, Mistral, and DeepSeek at Q3 through FP16. Lookup tables for 7B, 9B, 13B, 27B, 32B, 70B, and 120B models with real measurements and GPU recommendations. Updated July 2026.
- Ollama vs LM Studio: Speed, Setup, and Verdict Ollama gives you a CLI with 100+ models and an OpenAI-compatible API. LM Studio gives you a visual GUI with one-click downloads. Most power users run both—here's when to use each.
Getting Started (8)
- Ollama Troubleshooting Guide: Every Common Problem and Fix GPU not detected? Running at 1/30th speed on CPU? OOM crashes mid-generation? Every common Ollama error with exact diagnostic commands and fixes for Mac, Windows, and Linux. Updated July 2026 for v0.31.x and Qwen 3.5 + 3.6.
- Run Your First Local LLM in 15 Minutes Install Ollama, pull a model, and chat with AI offline—all in 15 minutes. Works on any Mac, Windows, or Linux machine with 8GB RAM. No accounts, no API keys, no fees.
- Ollama 0.30.0: What's New, What's Faster, What Breaks on Upgrade Ollama 0.30.0: llama.cpp integration, flash-attention default for Qwen/Gemma, broader model support. Firsthand upgrade notes, known issues to watch.
- Ubuntu 26.04 Is Built for Local AI — What Actually Changes Ubuntu 26.04 LTS packages NVIDIA CUDA and AMD ROCm in official repos. No more external downloads or dependency nightmares. What's confirmed and what it means for local AI.
- Qwen 3.5 Locally — 27B vs 35B-A3B vs 122B, Which Model Fits Your GPU Qwen 3.5 and 3.6 on local hardware. 27B dense vs 35B-A3B MoE vs 122B compared. VRAM tables, community tok/s on RTX 3090, and which to pick for your card.
Hardware & GPUs (54)
- Best Local LLMs for Mac in 2026 — M1 through M5 Tested Best model for every Mac tier, 8GB to the 96GB Studio ceiling. Qwen 3.6, Llama 4 Scout, DeepSeek V4, MLX vs Ollama. Why bandwidth, not RAM, sets your speed.
- GPU Buying Guide for Local AI: Pick the Right Card The complete GPU buying guide for local AI. Covers RTX 3060 through 4090 with VRAM analysis, performance benchmarks, prices, and used vs new buying advice.
- GTX 1650 vs RTX 3060 on a 35B MoE: What the Card Buys A $60-class GTX 1650 4 GB runs Qwen3.6-35B-A3B at 20 tok/s on 32 GB of RAM. The RTX 3060 in the same slot does 28 at the same setting and 39 tuned. Measured.
- A 177B Model on a 3060: The 32 GB Number Nobody Measured Qwen3.8-Flash-Next on an RTX 3090 and a 3060. Stock, the 3090 matches the video's 3060. One flag buys 34 to 40 percent. A real 32 GB box: 8.7 tok/s, not 22.
- The $36 RAM Fix That Made CPU Inference 56% Faster Adding a second RAM stick to a mini PC lifted CPU token generation 52-58% across four models. Prompt processing moved under 2%. Measured before and after.
Mac & Apple Silicon (14)
- Best Local LLMs for Mac in 2026 — M1 through M5 Tested Best model for every Mac tier, 8GB to the 96GB Studio ceiling. Qwen 3.6, Llama 4 Scout, DeepSeek V4, MLX vs Ollama. Why bandwidth, not RAM, sets your speed.
- Best Apple M5 Pro and Max for Local AI (2026) M5 Pro at 307GB/s, M5 Max at 614GB/s (or 460 on the 32-core bin), up to 128GB — now the highest-memory Mac you can buy. Picks for Qwen 3.6 and Llama 3.3 70B.
- OpenClaw on Mac: Setup, Optimization, and What Actually Works brew install openclaw-cli, connect Ollama, configure the gateway, and stop fighting macOS. Apple Silicon setup, memory math, launchd config, and the gotchas nobody warns you about.
- What Can You Run on 8GB Apple Silicon? Local AI on a Budget Mac Llama 3.2 3B runs at 30 tok/s. Phi-4 Mini fits with room to spare. 7B models technically load but swap to disk. Honest benchmarks and real limits for 8GB M1/M2/M3/M4 Macs.
- Stable Diffusion on Mac: Image Generation with MLX and Draw Things Draw Things generates SD 1.5 images in 8-15 seconds on an M2 Pro. ComfyUI takes 3x longer. MLX is fastest but code-only. Complete Mac image gen guide with speed tests.
Image Generation (11)
- Stable Diffusion Locally: Getting Started SD 1.5 runs on 4GB VRAM, SDXL needs 8GB, Flux needs 12GB+. Generate unlimited images for free in under 5 minutes with Fooocus or ComfyUI. Setup, models, and first image tips.
- Local AI Upscaling: Make Blurry Images Sharp Without the Cloud Upscayl, Real-ESRGAN, chaiNNer, and ComfyUI can upscale your photos for free on your own hardware. No subscriptions, no uploads, no per-image fees. Even a GTX 1060 works. Here's how to pick the right tool and start.
- Best Photorealism Checkpoints for Local Image Generation (2026) Juggernaut XL, RealVisXL, Realistic Vision, and Flux compared for photorealistic AI images. VRAM requirements, recommended settings, sample prompts, and installation for ComfyUI and A1111.
- Stable Diffusion on Mac: Image Generation with MLX and Draw Things Draw Things generates SD 1.5 images in 8-15 seconds on an M2 Pro. ComfyUI takes 3x longer. MLX is fastest but code-only. Complete Mac image gen guide with speed tests.
- SDXL vs SD 1.5 vs Flux: Which Image Model Should You Run Locally? SDXL vs SD 1.5 vs Flux compared by VRAM, speed, and quality. SD 1.5 needs 4GB, SDXL needs 8GB, Flux needs 12GB+. Benchmarks on real GPUs inside.
Models (47)
- Qwen 3.6 Complete Guide: 27B Dense, 35B-A3B MoE, and Which to Use Qwen 3.6 landed in two open-weight flavors: 27B dense and 35B-A3B MoE. Benchmarks, hardware fit, and which variant to run on your GPU.
- Best Qwen 3.5 Setup: When to Stay vs Move to 3.6 (2026) 3.5 is Qwen's stable open workhorse — 3.6 replaced only two tiers, 3.7 went closed. Which 3.5 model on which GPU, when to stay vs move to 3.6.
- Quantization Explained: What It Means for Local AI Q4_K_M shrinks a 7B model from 14GB to ~4GB while keeping 90-95% quality. What every quantization format means, how much VRAM each saves, and which to pick for your GPU.
- Who Actually Built Your Open Model? Soofi S vs Trinity Two independent labs shipped competitive open MoE models in 2026. I read both technical reports. One of them is built on NVIDIA's architecture, data and tokenizer.
- Qwen 3.6 MoE Routing, Measured: Flat Is the Wrong Number I traced every expert routing decision Qwen 3.6-35B-A3B makes across six workloads on an RTX 3060. Routing isn't flat, and 112 slots is the whole answer.
Software & Tools (29)
- Local LLMs vs ChatGPT: An Honest Comparison ChatGPT has web search, voice mode, and GPT-5.2. Local LLMs have privacy, no subscriptions, and no rate limits. Here's when each one wins, what the cost math actually looks like, and why most power users run both.
- llama.cpp vs Ollama vs vLLM: One User vs Many (2026) Single-user, the three are closer than benchmark posts admit. Concurrent, vLLM pulls 10-20x ahead. Decision tree, the vLLM VRAM gotcha, mid-2026 versions.
- Run Your First Local LLM in 15 Minutes Install Ollama, pull a model, and chat with AI offline—all in 15 minutes. Works on any Mac, Windows, or Linux machine with 8GB RAM. No accounts, no API keys, no fees.
- Every SSD-Streaming MoE Engine: What's Real, What's Dead Eight engines that stream MoE experts from disk appeared in four months. Three have no license file at all. Here's the verified state of each.
- How to Fix Slow Qwen 3.6 27B on RTX 3090 (10-80 tok/s) Qwen 3.6-27B at 12 tok/s on a 3090 when others report 35? The 8-step diagnostic checklist for offload, quants, templates, power limits, and backend choice.
AI Agents & OpenClaw (49)
- LoRA Skill Compilation Is a Double-Headed Coin Flip Ten LoRA seeds on identical data spread 3.62 points, against 4.65 for prompt compilation. Not one beat the no-adapter base. Measured on an RTX 3090.
- Skills in the Weights: The LoRA Answer to the Prompt Tax Compiling a skill into a prompt costs 1,383.9 tokens per call, forever. Putting it in a LoRA costs 7.55 GB and a training run. Only one has been measured.
- I Got the AI Result I Wanted. Then I Ran It Nine More Times A frontier model read my logs and wrote a skill that beat my local 27B baseline by 10.6 points. Nine more compilation runs showed the number was fake.
- Qwen 3.6: Why Q4 Quant Breaks Local Coding Agents (And the Fix) A viral thread says Q4-to-Q6 fixes Qwen 3.6 coding, but the test was confounded. What four independent reports show about the quant tax on coding agents.
- Anthropic Just Cut Off OpenClaw Users — Why Local Models Matter More Than Ever Starting April 4, Claude subscribers can no longer use their subscription for OpenClaw and other third-party harnesses. If you were relying on cloud AI for your agent, here's how to go fully local.
Use Cases (41)
- Local LLMs vs ChatGPT: An Honest Comparison ChatGPT has web search, voice mode, and GPT-5.2. Local LLMs have privacy, no subscriptions, and no rate limits. Here's when each one wins, what the cost math actually looks like, and why most power users run both.
- Stable Diffusion Locally: Getting Started SD 1.5 runs on 4GB VRAM, SDXL needs 8GB, Flux needs 12GB+. Generate unlimited images for free in under 5 minutes with Fooocus or ComfyUI. Setup, models, and first image tips.
- Local AI for Accounting and Tax: Keep Your Financial Data Off the Cloud Local LLMs can categorize transactions, draft client letters, extract receipt data, and answer questions over tax documents — without sending a single number to OpenAI or Google. What works, what doesn't, and how to set it up.
- Running OpenClaw on 4GB, 6GB, and 8GB GPUs: What Actually Works OpenClaw on low VRAM GPUs: 4GB is rough, 6GB is marginal, 8GB is where it starts working. Model picks, quantization tricks, partial offload, and when to just use a cloud API instead.
- Local AI for Therapists: Session Notes, Treatment Plans, and Client Privacy Without the Cloud Run AI on your own hardware to draft session notes, treatment plans, and clinical letters without sending client data to OpenAI. HIPAA-friendly setup for therapists.
Architecture & Theory (18)
- Quantization Explained: What It Means for Local AI Q4_K_M shrinks a 7B model from 14GB to ~4GB while keeping 90-95% quality. What every quantization format means, how much VRAM each saves, and which to pick for your GPU.
- Nvidia PAIR Bought Me 7 Percent. My Own Router Cost Me Half. Nvidia's new home-network AI router split 20 requests across an RTX 3090 and a 3060 for 1.07x over the 3090, then dropped half a 27B queue. Mine: 0.52x.
- Skills in the Weights: The LoRA Answer to the Prompt Tax Compiling a skill into a prompt costs 1,383.9 tokens per call, forever. Putting it in a LoRA costs 7.55 GB and a training run. Only one has been measured.
- 12 Architecture Patterns from the Claude Code Leak -- Ranked by Payoff for Local AI Claude Code's leaked source reveals 12 engineering patterns that power a $2.5B product. Ranked by how much each one improves your local AI agent setup.
- TurboQuant Explained: How Google's KV Cache Trick Cuts Memory 6x With Zero Quality Loss Google's TurboQuant compresses the KV cache 6x with zero accuracy loss. Here's what it actually does, how it works in llama.cpp and MLX, and what it means for running bigger models on your GPU.
Troubleshooting (21)
- Ollama Troubleshooting Guide: Every Common Problem and Fix GPU not detected? Running at 1/30th speed on CPU? OOM crashes mid-generation? Every common Ollama error with exact diagnostic commands and fixes for Mac, Windows, and Linux. Updated July 2026 for v0.31.x and Qwen 3.5 + 3.6.
- Why Qwen 3.8 27B Feels Slow: Reasoning Tokens Measured Qwen 3.8 27B generates at full speed on a 3090 and still crawls. Four runs, two models, two seeds: 92.8% of output is thinking, and the spread runs 320x to 542x.
- LLM Running Slow? Two Different Problems, Two Different Fixes Slow local LLM? Separate time-to-first-token from generation speed. Fix prompt processing with batch size and Flash Attention. Fix tok/s with GPU layers, quantization, and context length.
- Why Your Local LLM Is Slow: The num_ctx VRAM Overflow Nobody Warns You About DeepSeek-R1 14B went from 35 tok/s to 4.8 tok/s on the same GPU. The fix was one parameter. How num_ctx silently overflows VRAM and kills inference speed.
- Qwen2.5-VL Not Loading in LM Studio? Fix mmproj and Vision Errors Fix every Qwen2.5-VL error in LM Studio: missing mmproj, 'model type not supported', no eye icon, vision crashes. Exact fixes with file paths.