Guides
249 practical guides for running AI locally — from first install to advanced optimization.
Recently Updated
- Hugging Face Hacked by AI Agent — Saved by a Local Model (2026) Hugging Face says an autonomous AI agent breached its internal infra on July 16. The models you download are safe — here's what was and wasn't hit.
- Qwen 3.8 & Kimi K3: Open in Name, Closed in Practice — Run This Instead (2026) Qwen 3.8 (2.4T) and Kimi K3 (2.8T) both went 'open' in ten days. Neither fits your GPU. Here's Qwen's real open-weight cadence and what to run today.
- Best Uncensored Local LLMs by VRAM Tier (2026) Qwen 3.6 abliterated, Gemma 4 Heretic, Dolphin 3.0 — the current uncensored picks by VRAM tier. The Llama 3.1 / Qwen 2.5 era is mostly superseded. HuggingFace repos for every pick.
- Run LLMs on Mac M-Series: Faster, Without the Gotchas (2026) Foundational how-to for Apple Silicon local AI: unified memory, MLX vs Ollama vs llama.cpp Metal, verification, and the headless Mac Mini AI server.
- Free Local AI vs Paid Cloud APIs: Real Cost Comparison An RTX 3090 pays for itself in 2 weeks of moderate API usage. Full break-even math for local vs OpenAI, Anthropic, and Google APIs with current 2026 pricing.
- Running 70B Models Locally — Exact VRAM by Quantization Llama 3.3 70B needs 43GB at Q4, 75GB at Q8, 141GB at FP16. Every quant level, which GPUs fit, real speeds, and when a 27B or MoE model is the smarter buy.
- Mac Mini M4 for Local AI: Which Config to Buy and What It Actually Runs Mac Mini M4 Pro 48GB runs Qwen 3.6-35B-A3B silently at 40W. Which config to buy after Apple's 2026 price hikes, and what each tier actually runs for local AI.
- M4 Max and M3 Ultra for Local LLMs: Apple Silicon in 2026 No M4 Ultra exists. After Apple's 2026 memory cuts, the Mac Studio pairs the M4 Max (64GB) with the M3 Ultra (96GB, 819 GB/s). Which to buy for local AI.
- Best New Ollama 0.17 Features: ollama launch, MLX, and OpenClaw Support Everything new in Ollama 0.16 through 0.17.7: ollama launch for coding tools, native MLX on Apple Silicon, OpenClaw integration, web search API, and image generation. Updated March 2026.
- Mac Studio for Local AI: Is It Worth the Price? Mac Studio M4 Max (64GB) and M3 Ultra (96GB) for local LLMs after Apple's 2026 memory cuts. Real tok/s, cost vs dual RTX 3090, and who should buy one.
- Claude Code vs PI Agent — Which Coding Agent for Local AI? Claude Code vs PI Agent compared for local AI development. System prompts, tools, pricing, local model support, and honest verdicts for every type of developer.
- Best Apple M5 Pro and Max for Local AI (2026) M5 Pro at 307GB/s, M5 Max at 614GB/s (or 460 on the 32-core bin), up to 128GB — now the highest-memory Mac you can buy. Picks for Qwen 3.6 and Llama 3.3 70B.
- Best Local LLMs for Mac in 2026 — M1 through M5 Tested Best model for every Mac tier, 8GB to the 96GB Studio ceiling. Qwen 3.6, Llama 4 Scout, DeepSeek V4, MLX vs Ollama. Why bandwidth, not RAM, sets your speed.
- Inkling 975B vs Your 3090: The Real Memory Math (2026) Inkling's smallest 1-bit quant is 270GB. A maxed consumer desktop holds 256GB. The open frontier left consumer hardware behind. Here's the honest math.
- Llama 3 Guide: Every Size from 1B to 405B Complete Llama 3 guide covering every model from 1B to 405B. VRAM requirements, Ollama setup, benchmarks vs Qwen 3, and which size fits your hardware.
- DeepSeek Models Guide: R1, V3, and Coder Complete DeepSeek models guide covering R1, V3, and Coder locally. Which distilled R1 to pick for your GPU, VRAM requirements, and benchmarks vs Qwen 3.
- Best Vision Models You Can Run Locally: Every Model, Every GPU Tier Qwen 3.6 and Gemma 4 are the new local vision SOTA picks. Full VRAM table, Ollama commands, setup for every GPU from 4GB to 48GB+. Updated July 2026.
- Best Local LLMs for Function Calling: Qwen 3.6, Gemma 4 Function calling with local LLMs on Ollama and llama.cpp. Current lineup: Qwen 3.6, Gemma 4, DeepSeek V4. Common failures, agentic loop patterns. July 2026.
Most Popular
- Qwen 3.6 Complete Guide: 27B Dense, 35B-A3B MoE, and Which to Use Qwen 3.6 landed in two open-weight flavors: 27B dense and 35B-A3B MoE. Benchmarks, hardware fit, and which variant to run on your GPU.
- Best Qwen 3.5 Setup: When to Stay vs Move to 3.6 (2026) 3.5 is Qwen's stable open workhorse — 3.6 replaced only two tiers, 3.7 went closed. Which 3.5 model on which GPU, when to stay vs move to 3.6.
- Best Local LLMs for Mac in 2026 — M1 through M5 Tested Best model for every Mac tier, 8GB to the 96GB Studio ceiling. Qwen 3.6, Llama 4 Scout, DeepSeek V4, MLX vs Ollama. Why bandwidth, not RAM, sets your speed.
- llama.cpp vs Ollama vs vLLM: One User vs Many (2026) Single-user, the three are closer than benchmark posts admit. Concurrent, vLLM pulls 10-20x ahead. Decision tree, the vLLM VRAM gotcha, mid-2026 versions.
- Run Your First Local LLM in 15 Minutes Install Ollama, pull a model, and chat with AI offline—all in 15 minutes. Works on any Mac, Windows, or Linux machine with 8GB RAM. No accounts, no API keys, no fees.
- GPU Buying Guide for Local AI: Pick the Right Card The complete GPU buying guide for local AI. Covers RTX 3060 through 4090 with VRAM analysis, performance benchmarks, prices, and used vs new buying advice.
- Hugging Face Hacked by AI Agent — Saved by a Local Model (2026) Hugging Face says an autonomous AI agent breached its internal infra on July 16. The models you download are safe — here's what was and wasn't hit.
- Inkling 975B vs Your 3090: The Real Memory Math (2026) Inkling's smallest 1-bit quant is 270GB. A maxed consumer desktop holds 256GB. The open frontier left consumer hardware behind. Here's the honest math.
- China Made Open Source a Strategy. If It Pulls Back, Who Fills the Gap? If China restricts future open weights, who fills the gap? The West's open-model capacity is real (Ai2, Mistral, Apertus) but scattered and underfunded.
- China May Restrict Its AI Exports — Your Local Models Don't Care China's commerce ministry met Alibaba, ByteDance and Z.ai about walling off top AI models. The US did the same in June. The weights on your disk don't care.
- A100 vs H100 vs L40S vs 4090: Why the Cheaper GPU Costs More to Train On The cheapest GPU per hour is rarely the cheapest per training run. Real 2026 rental prices and total-cost math across the 4090, L40S, A100, H100, and H200.
- Best 24GB Backend Shootout: ik_llama vs BeeLlama vs llama.cpp ik_llama and BeeLlama both finish in 22-23s on the am17an 9-prompt harness vs mainline llama.cpp's 37s — 1.66x and 1.62x speedups via opposite strategies.
- Qwen 3.7 Open Weights Watch: The June Window Is Closing June 19: Qwen 3.7 open weights overdue. 3.5→3.6 cadence projected June 6-14; we're past. Max May 20 (56.6 AAI), VLA, Plus closed. Bench at drop.
- Wicked Fast Qwen 3.6 27B: 60 tok/s with MTP on RTX 3090 (2026) Firsthand bench: 60 tok/s on Qwen 3.6 27B Q4_K_M with MTP on a single RTX 3090 — 1.86x wall-clock speedup over baseline. PR #22673 progress May 6 → May 19.
- Wicked Fast Gemma 4 vs Qwen 3.6 on RTX 3090: 3.10x Tested Same RTX 3090, same llama.cpp build, same bench. Gemma 4 26B-A4B Q4_K_XL: 128 tok/s mean. Qwen 3.6-27B Q4_K_M: 41 tok/s. 3.10x faster, firsthand.
- DFlash vs MTP on RTX 3090: I Tested Both Locally Firsthand head-to-head bench of DFlash + DDTree against MTP (PR #22673) on a single RTX 3090, same Qwen 3.6-27B target. Real numbers, both backends.
- How to Fix Slow Qwen 3.6 27B on RTX 3090 (10-80 tok/s) Qwen 3.6-27B at 12 tok/s on a 3090 when others report 35? The 8-step diagnostic checklist for offload, quants, templates, power limits, and backend choice.
- Lightning 2.6.x Malware: Check Your Local AI Stack PyPI's lightning package was poisoned April 30 with malware that abuses Claude Code hooks. Here's the 5-minute audit I ran on my own 3090 box.
- How to Get 2.5x Faster Qwen on RTX 3090 (Free) I built DFlash on my RTX 3090 and ran the full bench. Real 2.5x speedup on Qwen 3.5 and 3.6 — below the 3.43x README claim, still huge. Here's how.
- Best Way to Run Qwen 3.6 35B MoE Locally: VRAM, Speed, Setup Qwen 3.6-35B-A3B has 35B total params but only 3B active per token. Real tok/s on RTX 3090, 4090, 5070 Ti, dual 5060 Ti, and M3 Ultra. Quants and setup.
- Best Way to Get 2x Token Output on RTX 3090: Qwen 3.6 + DFlash Luce DFlash + DDTree pushes Qwen 3.6-27B Q4_K_M from 35 tok/s to 69 tok/s on a single RTX 3090. Real benchmarks, setup, and honest limits.
- FP4 Just Landed in llama.cpp: NVFP4 vs MXFP4 Explained (2026) NVFP4 in llama.cpp, MXFP4 in ik_llama.cpp. The first practical FP4 quantization for the GGUF ecosystem — what works, what doesn't, and what to test.
- DeepSeek V4 Flash vs Pro: Verdict, Cost, Setup (2026) Flash or Pro? Which DeepSeek V4 to run, what each costs (Haiku-tier pricing), and how to deploy locally. The verdict, the tradeoffs, no launch recap.
- Anthropic Just Cut Off OpenClaw Users — Why Local Models Matter More Than Ever Starting April 4, Claude subscribers can no longer use their subscription for OpenClaw and other third-party harnesses. If you were relying on cloud AI for your agent, here's how to go fully local.
- OpenClaw Critical Sandbox Escape: Update to 2026.3.28 Now Ant AI Security Lab found 33 vulnerabilities in OpenClaw including critical privilege escalation and filesystem sandbox escape. If you're self-hosting, update immediately.
- TurboQuant Explained: How Google's KV Cache Trick Cuts Memory 6x With Zero Quality Loss Google's TurboQuant compresses the KV cache 6x with zero accuracy loss. Here's what it actually does, how it works in llama.cpp and MLX, and what it means for running bigger models on your GPU.
- Intel's $949 GPU Has 32GB VRAM and 608 GB/s Bandwidth: What It Means for Local AI Intel is launching a 32GB VRAM GPU for $949. Here's how it compares to the RTX 3090, RTX 4090, and used GPU market for running local LLMs and Stable Diffusion.
- Best Apple M5 Pro and Max for Local AI (2026) M5 Pro at 307GB/s, M5 Max at 614GB/s (or 460 on the 32-core bin), up to 128GB — now the highest-memory Mac you can buy. Picks for Qwen 3.6 and Llama 3.3 70B.
- Best 8GB GPU Model: How to Set Up Qwen 3.5 9B (Step by Step) Qwen 3.5 9B fits in 6.6GB and beats Qwen 3-class models 3x its size. Setup on Ollama/llama.cpp, quant table, where 9B still fits in the May 2026 lineup.
- Qwen 3.5 Small Models: The 9B Beats Last-Gen 30B — Here's What Matters for Local AI Alibaba's Qwen 3.5 drops 4 small models (0.8B to 9B) — all natively multimodal, 262K context, Apache 2.0. The 9B beats Qwen3-30B on reasoning and destroys GPT-5-Nano on vision. VRAM tables and what to run.
- Best Anime and Stylized Checkpoints for Local Image Generation (2026) Illustrious XL, NoobAI-XL, Animagine, Pony Diffusion, and SD 1.5 anime models compared. VRAM requirements, Danbooru prompting, LoRA picks, and settings for ComfyUI and A1111.
- Best Photorealism Checkpoints for Local Image Generation (2026) Juggernaut XL, RealVisXL, Realistic Vision, and Flux compared for photorealistic AI images. VRAM requirements, recommended settings, sample prompts, and installation for ComfyUI and A1111.
- Replace GitHub Copilot With Local LLMs in VS Code — Free, Private, No Subscription Set up free, private AI code completion in VS Code with Continue + Ollama. Autocomplete, chat, and agentic coding with Qwen models at every VRAM tier. Step-by-step setup, model picks, honest tradeoffs.
- Best Qwen 3.5 Models Ranked: Every Size, Every GPU, Every Quant Complete ranking of all Qwen 3.5 models from 0.8B to 397B. VRAM requirements, speed benchmarks, and which model to pick for your hardware.
- DeepSeek V4: Everything We Know Before It Drops DeepSeek V4 launches next week with native image and video generation, 1M context, and rumored 1T MoE params with only 32B active. Here's what local AI builders need to know and how to prepare.
- OpenClaw Security Report: February 2026 — ClawHub Malware, Google Suspensions, and Critical Fixes 17 security fixes, 341 malicious ClawHub skills, Google banning users, and the creator leaving. Every OpenClaw security event from February 2026.
- Best Local Alternatives to Claude Code in 2026 Aider, Continue.dev, Cline, OpenCode, Void, and Tabby compared. Which open-source coding tools work best with local models on your own GPU?
- Best OpenClaw Alternatives: 11 Tools That Actually Work in 2026 Tested alternatives to OpenClaw for local AI agent workflows — Hermes Agent, Pi Agent, MMX-CLI, VT Code, and seven more. Ranked by setup ease, model support, and what actually works after Anthropic's April 4 subscription cutoff.
- Best OpenClaw Tools and Extensions in 2026 Crabwalk visualizes agent actions in real time, Tokscale catches API bills before they hit $200+, and openclaw-docker locks down deployment. The best 3rd-party tools ranked.
- Fix OpenClaw Token Waste: $150 to $6 Overnight Cut OpenClaw API costs by 97% with three proven fixes: route heartbeats through Ollama, add tiered model routing, and purge session history token bloat.
- OpenClaw ClawHub Alert: 1,103 Malicious Skills Found OpenClaw ClawHub security alert: 1,103 malicious skills found across 14,706 audited. CVE-2026-28458 Browser Relay auth bypass. How to protect yourself now.
- Best Local Models for OpenClaw 2026: Qwen 3.6 + DeepSeek V4 Qwen 3.6-27B dense ties Sonnet 4.6 on agentic coding; 3.6-35B-A3B runs OpenClaw on 16GB VRAM. Plus DeepSeek V4-Flash, sampling tips, VRAM tiers.
- Run LLMs on Mac M-Series: Faster, Without the Gotchas (2026) Foundational how-to for Apple Silicon local AI: unified memory, MLX vs Ollama vs llama.cpp Metal, verification, and the headless Mac Mini AI server.
- Best Way to Set Up OpenClaw (2026 Guide) Run `npx openclaw@latest`, scan a QR code for WhatsApp, and your AI agent is live. Gateway needs just 2-4GB RAM. Add Ollama for local models or connect Claude/GPT-4 via API.
- Best Local Coding Models Ranked: Every VRAM Tier, Every Benchmark (2026) The best local LLMs for coding in 2026, ranked by VRAM tier. Qwen 3.6-27B, 3.6-35B-A3B, DeepSeek V4-Flash, benchmarks, editor setup, and Claude Code alternatives.
- Best VRAM Cheat Sheet for Local LLMs: Every Model, Every Quant Exact VRAM for Qwen 3.6, Qwen 3.5, Llama, Mistral, and DeepSeek at Q3 through FP16. Lookup tables for 7B, 9B, 13B, 27B, 32B, 70B, and 120B models with real measurements and GPU recommendations. Updated July 2026.
- Ollama vs LM Studio: Speed, Setup, and Verdict Ollama gives you a CLI with 100+ models and an OpenAI-compatible API. LM Studio gives you a visual GUI with one-click downloads. Most power users run both—here's when to use each.
Getting Started (8)
- Ollama Troubleshooting Guide: Every Common Problem and Fix GPU not detected? Running at 1/30th speed on CPU? OOM crashes mid-generation? Every common Ollama error with exact diagnostic commands and fixes for Mac, Windows, and Linux. Updated July 2026 for v0.31.x and Qwen 3.5 + 3.6.
- Run Your First Local LLM in 15 Minutes Install Ollama, pull a model, and chat with AI offline—all in 15 minutes. Works on any Mac, Windows, or Linux machine with 8GB RAM. No accounts, no API keys, no fees.
- Ollama 0.30.0: What's New, What's Faster, What Breaks on Upgrade Ollama 0.30.0: llama.cpp integration, flash-attention default for Qwen/Gemma, broader model support. Firsthand upgrade notes, known issues to watch.
- Ubuntu 26.04 Is Built for Local AI — What Actually Changes Ubuntu 26.04 LTS packages NVIDIA CUDA and AMD ROCm in official repos. No more external downloads or dependency nightmares. What's confirmed and what it means for local AI.
- Qwen 3.5 Locally — 27B vs 35B-A3B vs 122B, Which Model Fits Your GPU Qwen 3.5 and 3.6 on local hardware. 27B dense vs 35B-A3B MoE vs 122B compared. VRAM tables, community tok/s on RTX 3090, and which to pick for your card.
Hardware & GPUs (51)
- Best Local LLMs for Mac in 2026 — M1 through M5 Tested Best model for every Mac tier, 8GB to the 96GB Studio ceiling. Qwen 3.6, Llama 4 Scout, DeepSeek V4, MLX vs Ollama. Why bandwidth, not RAM, sets your speed.
- GPU Buying Guide for Local AI: Pick the Right Card The complete GPU buying guide for local AI. Covers RTX 3060 through 4090 with VRAM analysis, performance benchmarks, prices, and used vs new buying advice.
- AMD vs NVIDIA for Local AI: Is ROCm Finally Ready? ROCm 7.2 finally ships official RDNA4 support and one installer for Linux + Windows. The RX 7900 XTX (24GB) and new RX 9070 XT (16GB) are real options now. Honest mid-2026 benchmarks and the compatibility gaps that remain.
- A100 vs H100 vs L40S vs 4090: Why the Cheaper GPU Costs More to Train On The cheapest GPU per hour is rarely the cheapest per training run. Real 2026 rental prices and total-cost math across the 4090, L40S, A100, H100, and H200.
- How to Run GLM 5.2 Locally: GPU, VRAM & Quant Guide GLM 5.2 is 753B params and 1.51TB at full precision. Run it locally: the live Unsloth quant ladder, every GPU and RAM path, and the quant to actually target.
Mac & Apple Silicon (14)
- Best Local LLMs for Mac in 2026 — M1 through M5 Tested Best model for every Mac tier, 8GB to the 96GB Studio ceiling. Qwen 3.6, Llama 4 Scout, DeepSeek V4, MLX vs Ollama. Why bandwidth, not RAM, sets your speed.
- Best Apple M5 Pro and Max for Local AI (2026) M5 Pro at 307GB/s, M5 Max at 614GB/s (or 460 on the 32-core bin), up to 128GB — now the highest-memory Mac you can buy. Picks for Qwen 3.6 and Llama 3.3 70B.
- OpenClaw on Mac: Setup, Optimization, and What Actually Works brew install openclaw-cli, connect Ollama, configure the gateway, and stop fighting macOS. Apple Silicon setup, memory math, launchd config, and the gotchas nobody warns you about.
- What Can You Run on 8GB Apple Silicon? Local AI on a Budget Mac Llama 3.2 3B runs at 30 tok/s. Phi-4 Mini fits with room to spare. 7B models technically load but swap to disk. Honest benchmarks and real limits for 8GB M1/M2/M3/M4 Macs.
- Stable Diffusion on Mac: Image Generation with MLX and Draw Things Draw Things generates SD 1.5 images in 8-15 seconds on an M2 Pro. ComfyUI takes 3x longer. MLX is fastest but code-only. Complete Mac image gen guide with speed tests.
Image Generation (11)
- Stable Diffusion Locally: Getting Started SD 1.5 runs on 4GB VRAM, SDXL needs 8GB, Flux needs 12GB+. Generate unlimited images for free in under 5 minutes with Fooocus or ComfyUI. Setup, models, and first image tips.
- Local AI Upscaling: Make Blurry Images Sharp Without the Cloud Upscayl, Real-ESRGAN, chaiNNer, and ComfyUI can upscale your photos for free on your own hardware. No subscriptions, no uploads, no per-image fees. Even a GTX 1060 works. Here's how to pick the right tool and start.
- Best Photorealism Checkpoints for Local Image Generation (2026) Juggernaut XL, RealVisXL, Realistic Vision, and Flux compared for photorealistic AI images. VRAM requirements, recommended settings, sample prompts, and installation for ComfyUI and A1111.
- Stable Diffusion on Mac: Image Generation with MLX and Draw Things Draw Things generates SD 1.5 images in 8-15 seconds on an M2 Pro. ComfyUI takes 3x longer. MLX is fastest but code-only. Complete Mac image gen guide with speed tests.
- SDXL vs SD 1.5 vs Flux: Which Image Model Should You Run Locally? SDXL vs SD 1.5 vs Flux compared by VRAM, speed, and quality. SD 1.5 needs 4GB, SDXL needs 8GB, Flux needs 12GB+. Benchmarks on real GPUs inside.
Models (45)
- Qwen 3.6 Complete Guide: 27B Dense, 35B-A3B MoE, and Which to Use Qwen 3.6 landed in two open-weight flavors: 27B dense and 35B-A3B MoE. Benchmarks, hardware fit, and which variant to run on your GPU.
- Best Qwen 3.5 Setup: When to Stay vs Move to 3.6 (2026) 3.5 is Qwen's stable open workhorse — 3.6 replaced only two tiers, 3.7 went closed. Which 3.5 model on which GPU, when to stay vs move to 3.6.
- Quantization Explained: What It Means for Local AI Q4_K_M shrinks a 7B model from 14GB to ~4GB while keeping 90-95% quality. What every quantization format means, how much VRAM each saves, and which to pick for your GPU.
- Hugging Face Hacked by AI Agent — Saved by a Local Model (2026) Hugging Face says an autonomous AI agent breached its internal infra on July 16. The models you download are safe — here's what was and wasn't hit.
- Qwen 3.8 & Kimi K3: Open in Name, Closed in Practice — Run This Instead (2026) Qwen 3.8 (2.4T) and Kimi K3 (2.8T) both went 'open' in ten days. Neither fits your GPU. Here's Qwen's real open-weight cadence and what to run today.
Software & Tools (28)
- Local LLMs vs ChatGPT: An Honest Comparison ChatGPT has web search, voice mode, and GPT-5.2. Local LLMs have privacy, no subscriptions, and no rate limits. Here's when each one wins, what the cost math actually looks like, and why most power users run both.
- llama.cpp vs Ollama vs vLLM: One User vs Many (2026) Single-user, the three are closer than benchmark posts admit. Concurrent, vLLM pulls 10-20x ahead. Decision tree, the vLLM VRAM gotcha, mid-2026 versions.
- Run Your First Local LLM in 15 Minutes Install Ollama, pull a model, and chat with AI offline—all in 15 minutes. Works on any Mac, Windows, or Linux machine with 8GB RAM. No accounts, no API keys, no fees.
- Text Generation WebUI Setup Guide (2026) Install oobabooga's TextGen (formerly text-generation-webui), load GGUF/EXL2/EXL3 models, and configure GPU offloading. Now with vision, tool-calling, and an Anthropic-compatible API. Covers the settings most guides skip.
- How to Fix Slow Qwen 3.6 27B on RTX 3090 (10-80 tok/s) Qwen 3.6-27B at 12 tok/s on a 3090 when others report 35? The 8-step diagnostic checklist for offload, quants, templates, power limits, and backend choice.
AI Agents & OpenClaw (46)
- Qwen 3.6: Why Q4 Quant Breaks Local Coding Agents (And the Fix) A viral thread says Q4-to-Q6 fixes Qwen 3.6 coding, but the test was confounded. What four independent reports show about the quant tax on coding agents.
- Anthropic Just Cut Off OpenClaw Users — Why Local Models Matter More Than Ever Starting April 4, Claude subscribers can no longer use their subscription for OpenClaw and other third-party harnesses. If you were relying on cloud AI for your agent, here's how to go fully local.
- 12 Architecture Patterns from the Claude Code Leak -- Ranked by Payoff for Local AI Claude Code's leaked source reveals 12 engineering patterns that power a $2.5B product. Ranked by how much each one improves your local AI agent setup.
- OpenClaw Critical Sandbox Escape: Update to 2026.3.28 Now Ant AI Security Lab found 33 vulnerabilities in OpenClaw including critical privilege escalation and filesystem sandbox escape. If you're self-hosting, update immediately.
- Claude Code's Source Just Leaked: What 500K Lines of TypeScript Reveal About AI Coding Agents Claude Code's full source was exposed via npm source maps. Here's what the leaked architecture reveals about multi-agent orchestration, and what it means for local AI agent builders.
Use Cases (41)
- Local LLMs vs ChatGPT: An Honest Comparison ChatGPT has web search, voice mode, and GPT-5.2. Local LLMs have privacy, no subscriptions, and no rate limits. Here's when each one wins, what the cost math actually looks like, and why most power users run both.
- Stable Diffusion Locally: Getting Started SD 1.5 runs on 4GB VRAM, SDXL needs 8GB, Flux needs 12GB+. Generate unlimited images for free in under 5 minutes with Fooocus or ComfyUI. Setup, models, and first image tips.
- Local AI for Privacy: What's Actually Private Running AI locally keeps prompts off corporate servers — but model downloads, telemetry, and VS Code extensions can still leak data. Here's what's genuinely private, what isn't, and how to close every gap.
- Fine-Tuning LLMs on Consumer Hardware: LoRA and QLoRA Guide Fine-tune a 7B model on 6-10GB VRAM with QLoRA and Unsloth (2-5x faster, 70% less memory). Only 200-500 examples needed — and in 2026 you can fine-tune a 30B MoE on a single 24GB card. Dataset prep through training on RTX 3060-5090.
- Best Local LLMs for Translation: What Actually Works NLLB handles 200 languages on 3GB VRAM. Qwen 3 and Gemma 3 rival DeepL for European pairs. Opus-MT runs at 300MB per direction. Which local translation model fits your hardware and language needs.
Architecture & Theory (16)
- Quantization Explained: What It Means for Local AI Q4_K_M shrinks a 7B model from 14GB to ~4GB while keeping 90-95% quality. What every quantization format means, how much VRAM each saves, and which to pick for your GPU.
- Fine-Tuning LLMs on Consumer Hardware: LoRA and QLoRA Guide Fine-tune a 7B model on 6-10GB VRAM with QLoRA and Unsloth (2-5x faster, 70% less memory). Only 200-500 examples needed — and in 2026 you can fine-tune a 30B MoE on a single 24GB card. Dataset prep through training on RTX 3060-5090.
- 12 Architecture Patterns from the Claude Code Leak -- Ranked by Payoff for Local AI Claude Code's leaked source reveals 12 engineering patterns that power a $2.5B product. Ranked by how much each one improves your local AI agent setup.
- TurboQuant Explained: How Google's KV Cache Trick Cuts Memory 6x With Zero Quality Loss Google's TurboQuant compresses the KV cache 6x with zero accuracy loss. Here's what it actually does, how it works in llama.cpp and MLX, and what it means for running bigger models on your GPU.
- Model Routing for Local AI — Stop Using One Model for Everything You're running one model for every task. That wastes VRAM, burns electricity, and gives worse results. Model routing sends each task to the right model at the right cost. Here's how to set it up.
Troubleshooting (20)
- Ollama Troubleshooting Guide: Every Common Problem and Fix GPU not detected? Running at 1/30th speed on CPU? OOM crashes mid-generation? Every common Ollama error with exact diagnostic commands and fixes for Mac, Windows, and Linux. Updated July 2026 for v0.31.x and Qwen 3.5 + 3.6.
- LLM Running Slow? Two Different Problems, Two Different Fixes Slow local LLM? Separate time-to-first-token from generation speed. Fix prompt processing with batch size and Flash Attention. Fix tok/s with GPU layers, quantization, and context length.
- Why Your Local LLM Is Slow: The num_ctx VRAM Overflow Nobody Warns You About DeepSeek-R1 14B went from 35 tok/s to 4.8 tok/s on the same GPU. The fix was one parameter. How num_ctx silently overflows VRAM and kills inference speed.
- Qwen2.5-VL Not Loading in LM Studio? Fix mmproj and Vision Errors Fix every Qwen2.5-VL error in LM Studio: missing mmproj, 'model type not supported', no eye icon, vision crashes. Exact fixes with file paths.
- Open WebUI Not Connecting to Ollama? Every Fix Docker networking, wrong OLLAMA_BASE_URL, localhost confusion, WSL2 isolation, missing models, random disconnects. Every Open WebUI + Ollama connection problem with the exact fix.