# InsiderLLM - Local AI Guides > Practical guides for running AI locally on consumer hardware. Budget-focused, no fluff. InsiderLLM helps hobbyists and developers run large language models and image generators on their own hardware. We focus on what actually works with the GPUs and computers people already own. ## Data [Local LLM Benchmarks: Open Dataset]: /benchmarks/ - Firsthand local inference benchmarks measured on hardware we own: RTX 3090 and RTX 3060, Qwen 3.5 and 3.6, llama.cpp / ik_llama / BeeLlama. Full rig specs, exact flags, and per-prompt breakdowns. [Benchmark dataset (raw JSON)]: /benchmarks.json - The same dataset as machine-readable JSON. CC BY 4.0, free to reuse with attribution. 134 measured rows, each carrying hardware, model, config, result, and provenance back to a published article. ## Guides [12 Architecture Patterns from the Claude Code Leak -- Ranked by Payoff for Local AI]: /guides/claude-code-architecture-lessons-local-ai/ - Claude Code's leaked source reveals 12 engineering patterns that power a $2.5B product. Ranked by how much each one improves your local AI agent setup. [A 177B Model on a 3060: The 32 GB Number Nobody Measured]: /guides/qwen3-8-flash-next-rtx-3090-3060-32gb/ - Qwen3.8-Flash-Next on an RTX 3090 and a 3060. Stock, the 3090 matches the video's 3060. One flag buys 34 to 40 percent. A real 32 GB box: 8.7 tok/s, not 22. [A100 vs H100 vs L40S vs 4090: Why the Cheaper GPU Costs More to Train On]: /guides/cheaper-gpu-costs-more-training/ - The cheapest GPU per hour is rarely the cheapest per training run. Real 2026 rental prices and total-cost math across the 4090, L40S, A100, H100, and H200. [AI Agents Rebuilt Their Own Message Board in Two Days (2026)]: /guides/ai-agent-coordination-incidents-2026/ - Five reports now. OpenAI's technical report explains why agents built a message board: 198 unsolvable tasks and too much time. What transfers to a home cluster and what doesn't. [AI Art Styles & Workflows: SD and Flux Guide]: /guides/ai-art-styles-workflows-guide/ - Photorealism, anime, oil painting, concept art, and pixel art on 8GB+ VRAM. Model picks, LoRA stacking at 0.5-0.8 weight, and ComfyUI workflows for each style. [AI Kill Switch Act vs Open Weights: Can You Shut Down a File?]: /guides/ai-pacing-open-weights/ - 1,238 AI staff asked Washington to slow the frontier. Every mechanism proposed works at release — the only moment an open release can be touched. [AI Tool Sprawl: You're Running 6 AI Tools and None of Them Talk to Each Other]: /guides/ai-tool-sprawl-consolidation-guide/ - Ollama for local chat, LM Studio for testing, ChatGPT for the hard stuff, Claude for writing, Copilot in your editor, Open WebUI as a frontend. Six tools, zero integration. Here's how to consolidate without losing capability. [AMD vs NVIDIA for Local AI: Is ROCm Finally Ready?]: /guides/amd-vs-nvidia-local-ai-rocm/ - ROCm 7.2 finally ships official RDNA4 support and one installer for Linux + Windows. The RX 7900 XTX (24GB) and new RX 9070 XT (16GB) are real options now. Honest mid-2026 benchmarks and the compatibility gaps that remain. [Anthropic Just Cut Off OpenClaw Users — Why Local Models Matter More Than Ever]: /guides/anthropic-cuts-openclaw-claude-subscription/ - Starting April 4, Claude subscribers can no longer use their subscription for OpenClaw and other third-party harnesses. If you were relying on cloud AI for your agent, here's how to go fully local. [AnythingLLM Setup Guide: Chat With Your Documents Locally]: /guides/anythingllm-setup-guide/ - Upload PDFs, paste URLs, and chat with your files — no coding, no cloud. AnythingLLM connects to Ollama in 5 minutes with point-and-click RAG on 54K+ GitHub stars. [Apple Neural Engine for LLM Inference: What Actually Works]: /guides/apple-neural-engine-llm-inference/ - Apple Silicon has a dedicated Neural Engine that most LLM tools ignore. Here's what it can do for inference, what it can't, and whether ANE-based tools like ANEMLL are worth trying today. [Are Mistral Models Still Worth Running? Only Nemo 12B (Here's Why)]: /guides/mistral-mixtral-guide/ - Mistral Medium 3.5-128B dropped April 29, 2026: dense 128B, 256k context, Modified MIT. Hardware reality, license caveats, which Mistral to actually run. [Best 24GB Backend Shootout: ik_llama vs BeeLlama vs llama.cpp]: /guides/best-24gb-backend-shootout-ik-llama-beellama-llamacpp/ - ik_llama and BeeLlama both finish in 22-23s on the am17an 9-prompt harness vs mainline llama.cpp's 37s — 1.66x and 1.62x speedups via opposite strategies. [Best 8GB GPU Model: How to Set Up Qwen 3.5 9B (Step by Step)]: /guides/qwen-3-5-9b-setup-guide/ - Qwen 3.5 9B fits in 6.6GB and beats Qwen 3-class models 3x its size. Setup on Ollama/llama.cpp, quant table, where 9B still fits in the May 2026 lineup. [Best Anime and Stylized Checkpoints for Local Image Generation (2026)]: /guides/best-anime-stylized-checkpoints-local-image-generation/ - Illustrious XL, NoobAI-XL, Animagine, Pony Diffusion, and SD 1.5 anime models compared. VRAM requirements, Danbooru prompting, LoRA picks, and settings for ComfyUI and A1111. [Best Apple M5 Pro and Max for Local AI (2026)]: /guides/apple-m5-pro-max-local-ai/ - M5 Pro at 307GB/s, M5 Max at 614GB/s (or 460 on the 32-core bin), up to 128GB — now the highest-memory Mac you can buy. Picks for Qwen 3.6 and Llama 3.3 70B. [Best Docker Setup for Local AI: Ollama + Open WebUI (2026)]: /guides/docker-local-ai-ollama-open-webui-gpu-passthrough/ - Five copy-paste Docker compose recipes for Ollama 0.24 + Open WebUI: NVIDIA Blackwell, AMD ROCm, multi-GPU, and CPU. Plus the Apple Silicon catch most guides skip. [Best Dual-GPU Local AI Setup: RTX 3090, 5060 Ti (2026)]: /guides/multi-gpu-local-ai/ - Dual RTX 3090, 2x RTX 5060 Ti, 2x 2080 Ti modded, mixed setups: real configs for Qwen 3.6, MoE, 70B. Tensor vs pipeline parallelism, llama.cpp/vLLM. [Best GPU Under $300 for Local AI (2026 Picks)]: /guides/best-gpu-under-300-local-ai/ - Find the best GPU under $300 for local AI. We compare the RTX 3060 12GB, RX 7600, and Intel Arc B580 with VRAM analysis, LLM benchmarks, and real pricing. [Best GPU Under $500 for Local AI (2026 Picks)]: /guides/best-gpu-under-500-local-ai/ - Find the best GPU under $500 for running local AI in 2026. RTX 4060 Ti 16GB, used RTX 3080, RTX 3060 12GB, and RX 7700 XT compared with real benchmarks. [Best Hardware for Running OpenClaw — Mac Mini vs VPS vs Your Old PC]: /guides/openclaw-hardware-mac-mini-vps-pc/ - OpenClaw runs 24/7. A Mac Mini M4 draws 4 watts idle. A free Oracle VPS costs nothing. A used ThinkCentre costs $85. Here's which one to pick. [Best LLM Speed Trick: ExLlamaV2 vs llama.cpp Benchmarks (50-85% Faster)]: /guides/exllamav2-vs-llamacpp-speed-comparison/ - Head-to-head speed benchmarks on RTX 3090 and 4090. ExLlamaV2 generates tokens 50-85% faster than llama.cpp on NVIDIA GPUs. Full comparison with setup guides for both. [Best Local AI Video: Wan vs LTX-2.3 & More (2026)]: /guides/local-ai-video-generation/ - Wan leads on quality, LTX-2.3 renders in seconds with native audio, and quantization drops the floor to 6-8GB VRAM. Benchmarks and setup for consumer GPUs. [Best Local Alternatives to Claude Code in 2026]: /guides/local-alternatives-claude-code-2026/ - Aider, Continue.dev, Cline, OpenCode, Void, and Tabby compared. Which open-source coding tools work best with local models on your own GPU? [Best Local Coding Models Ranked: Every VRAM Tier, Every Benchmark (2026)]: /guides/best-local-coding-models-2026/ - The best local LLMs for coding in 2026, ranked by VRAM tier. Qwen 3.6-27B, 3.6-35B-A3B, DeepSeek V4-Flash, benchmarks, editor setup, and Claude Code alternatives. [Best Local FLUX Setup: FLUX.2, FLUX.1, RTX 3090 (2026)]: /guides/flux-locally-complete-guide/ - FLUX.1 vs FLUX.2: 12GB VRAM minimum with GGUF, 24GB at full quality. ComfyUI / Forge / Python setup, May 2026 hardware, NVFP4 on Blackwell. [Best Local LLMs for Chat & Conversation]: /guides/best-local-llms-chat-conversation/ - The best local LLMs for chat and conversation in 2026. Picks for every VRAM tier from 8GB to 24GB, with Ollama commands to start chatting immediately. [Best Local LLMs for Data Analysis (2026)]: /guides/best-local-llms-data-analysis/ - Which local models write the best pandas and SQL code on your own hardware. Tested Qwen 2.5 Coder, DeepSeek, and Llama on real datasets with accuracy scores. [Best Local LLMs for Function Calling: Qwen 3.6, Gemma 4]: /guides/function-calling-local-llms/ - Function calling with local LLMs on Ollama and llama.cpp. Current lineup: Qwen 3.6, Gemma 4, DeepSeek V4. Common failures, agentic loop patterns. July 2026. [Best Local LLMs for Mac in 2026 — M1 through M5 Tested]: /guides/best-local-llms-mac-2026/ - Best model for every Mac tier, 8GB to the 96GB Studio ceiling. Qwen 3.6, Llama 4 Scout, DeepSeek V4, MLX vs Ollama. Why bandwidth, not RAM, sets your speed. [Best Local LLMs for Math & Reasoning: What Actually Works]: /guides/best-local-llms-math-reasoning/ - The best local LLMs for math and reasoning in 2026, ranked by VRAM tier. AIME 2026 and GPQA benchmarks for Qwen 3.6, Qwen 3.5 thinking, and where the old R1-distills now stand. [Best Local LLMs for RAG in 2026]: /guides/best-local-llms-rag/ - The best local models for retrieval-augmented generation by VRAM tier. Qwen 3, Command R 35B, and the current embedding picks (EmbeddingGemma, bge-m3, Qwen3-Embedding) — with the real failure modes. [Best Local LLMs for Structured Output: Qwen 3.6, Gemma 4]: /guides/structured-output-local-llms/ - JSON schema, grammar constraints, and Outlines compared. Current model picks: Qwen 3.6, Gemma 4, DeepSeek V4. Common failures + working code. May 2026. [Best Local LLMs for Summarization]: /guides/best-local-llms-summarization/ - Qwen 2.5 14B is the summarization sweet spot — strong instruction following, 128K context for 200-page docs, fits on 16GB VRAM. Model picks by use case, quality ratings, chunking strategies, and prompting tips. [Best Local LLMs for Translation: What Actually Works]: /guides/best-local-llms-translation/ - NLLB handles 200 languages on 3GB VRAM. Qwen 3 and Gemma 3 rival DeepL for European pairs. Opus-MT runs at 300MB per direction. Which local translation model fits your hardware and language needs. [Best Local LLMs for Writing & Creative Work]: /guides/best-local-llms-writing-creative-work/ - Llama 3.3 70B is the best local prose model in 2026; Qwen3 32B is the 24GB sweet spot for fiction and long-form. Model picks for every VRAM tier and writing task. [Best Local Models for OpenClaw 2026: Qwen 3.6 + DeepSeek V4]: /guides/best-local-models-openclaw/ - Qwen 3.6-27B dense ties Sonnet 4.6 on agentic coding; 3.6-35B-A3B runs OpenClaw on 16GB VRAM. Plus DeepSeek V4-Flash, sampling tips, VRAM tiers. [Best Local Models for PI Agent: Qwen 3.6, Gemma 4 (2026 Setup)]: /guides/pi-agent-local-models-ollama/ - PI Agent runs any model locally via Ollama. May 2026 picks: Qwen 3.6 27B / 35B-A3B MoE, Gemma 4 26B-A4B. Setup, model comparisons, honest limits. [Best Mini PCs for Local AI Under $300 in 2026]: /guides/best-mini-pcs-local-ai-2026/ - A $200 refurbished ThinkCentre runs 7B models at 5-8 tok/s. A $350 AMD Ryzen box hits 10-15 tok/s. Specific picks, real benchmarks, and what's worth buying. [Best Models Under 3B: Small LLMs That Work]: /guides/best-models-under-3b-parameters/ - The best models under 3B parameters for laptops, old GPUs, Raspberry Pi, and phones. What works, what doesn't, and which tiny LLM to pick for your use case. [Best New Ollama 0.17 Features: ollama launch, MLX, and OpenClaw Support]: /guides/ollama-0-17-new-features/ - Everything new in Ollama 0.16 through 0.17.7: ollama launch for coding tools, native MLX on Apple Silicon, OpenClaw integration, web search API, and image generation. Updated March 2026. [Best OpenClaw Alternatives: 11 Tools That Actually Work in 2026]: /guides/best-openclaw-alternatives/ - Tested alternatives to OpenClaw for local AI agent workflows — Hermes Agent, Pi Agent, MMX-CLI, VT Code, and seven more. Ranked by setup ease, model support, and what actually works after Anthropic's April 4 subscription cutoff. [Best OpenClaw Plugins and Skills Guide (2026)]: /guides/openclaw-plugins-skills-guide/ - Every OpenClaw skill worth installing, how to avoid malicious plugins on ClawHub, and how to build your own. 1,103 of 14,706 skills are malicious. [Best OpenClaw Tools and Extensions in 2026]: /guides/best-openclaw-tools-extensions/ - Crabwalk visualizes agent actions in real time, Tokscale catches API bills before they hit $200+, and openclaw-docker locks down deployment. The best 3rd-party tools ranked. [Best Photorealism Checkpoints for Local Image Generation (2026)]: /guides/best-photorealism-checkpoints-local-image-generation/ - Juggernaut XL, RealVisXL, Realistic Vision, and Flux compared for photorealistic AI images. VRAM requirements, recommended settings, sample prompts, and installation for ComfyUI and A1111. [Best Qwen 3.5 Models Ranked: Every Size, Every GPU, Every Quant]: /guides/qwen-3-5-local-guide/ - Complete ranking of all Qwen 3.5 models from 0.8B to 397B. VRAM requirements, speed benchmarks, and which model to pick for your hardware. [Best Qwen 3.5 Setup: When to Stay vs Move to 3.6 (2026)]: /guides/qwen-3-5-local-ai-guide/ - 3.5 is Qwen's stable open workhorse — 3.6 replaced only two tiers, 3.7 went closed. Which 3.5 model on which GPU, when to stay vs move to 3.6. [Best Qwen Models Ranked: Which to Run Locally (Mid-2026)]: /guides/qwen-models-guide/ - Every Qwen you can run locally: 3.8-27B (new, Apache 2.0), 3.6, 3.5, Qwen3-Coder-Next, Qwen-VL. VRAM per tier, Ollama setup, and how they bench. [Best Uncensored Local LLMs by VRAM Tier (2026)]: /guides/best-uncensored-local-llms/ - Qwen 3.6 abliterated, Gemma 4 Heretic, Dolphin 3.0 — the current uncensored picks by VRAM tier. The Llama 3.1 / Qwen 2.5 era is mostly superseded. HuggingFace repos for every pick. [Best Used GPUs for Local AI: 2026 Buying Guide]: /guides/best-used-gpus-local-ai-2026/ - RTX 3090 at ~$1,200-1,400 for 24GB, RTX 3060 12GB at $200-400, RTX 3080 at $350-400. Tier rankings, fair prices, what to avoid (skip the 8GB 3070), and where to buy safely. [Best Vision Models You Can Run Locally: Every Model, Every GPU Tier]: /guides/vision-models-locally/ - Qwen 3.6 and Gemma 4 are the new local vision SOTA picks. Full VRAM table, Ollama commands, setup for every GPU from 4GB to 48GB+. Updated July 2026. [Best VRAM Cheat Sheet for Local LLMs: Every Model, Every Quant]: /guides/vram-requirements-local-llms/ - Exact VRAM for Qwen 3.6, Qwen 3.5, Llama, Mistral, and DeepSeek at Q3 through FP16. Lookup tables for 7B, 9B, 13B, 27B, 32B, 70B, and 120B models with real measurements and GPU recommendations. Updated July 2026. [Best Way to Get 2x Token Output on RTX 3090: Qwen 3.6 + DFlash]: /guides/best-way-2x-token-output-rtx-3090-qwen-3-6-dflash/ - Luce DFlash + DDTree pushes Qwen 3.6-27B Q4_K_M from 35 tok/s to 69 tok/s on a single RTX 3090. Real benchmarks, setup, and honest limits. [Best Way to Run 31B Models on a Laptop? Treat Them Like Databases]: /guides/run-31b-models-laptop-larql/ - LARQL decompiles transformer weights into a queryable graph called a vindex. The project pitches a new shape for local inference: walk a subgraph, patch facts, stream from disk. Here's what's real, what's claimed, and what's still research. [Best Way to Run Qwen 3.5 on Mac: MLX vs Ollama Speed Test]: /guides/qwen35-mac-mlx-vs-ollama/ - MLX runs Qwen 3.5 up to 2x faster than Ollama on Apple Silicon. Head-to-head benchmarks on M1 through M4, with setup instructions for both. [Best Way to Run Qwen 3.6 35B MoE Locally: VRAM, Speed, Setup]: /guides/best-way-run-qwen-3-6-35b-moe-locally/ - Qwen 3.6-35B-A3B has 35B total params but only 3B active per token. Real tok/s on RTX 3090, 4090, 5070 Ti, dual 5060 Ti, and M3 Ultra. Quants and setup. [Best Way to Set Up OpenClaw (2026 Guide)]: /guides/openclaw-setup-guide/ - Run `npx openclaw@latest`, scan a QR code for WhatsApp, and your AI agent is live. Gateway needs just 2-4GB RAM. Add Ollama for local models or connect Claude/GPT-4 via API. [Best Ways to Connect Local AI to Notion in 2026]: /guides/notion-local-ai-integration/ - 4 real ways to connect Notion to a local LLM without sending data to the cloud. MCP servers, RAG pipelines, Open WebUI, and n8n workflows compared with setup steps. [Best Ways to Fix OpenClaw Tool Call Failures: 2026 Guide]: /guides/openclaw-tool-call-failures/ - Your OpenClaw agent silently fails, loops, or corrupts its session. Six debug paths plus May 2026 gotchas: Qwen 3.6 whitespace kwargs, Gemma 4 thinking mode. [Best Ways to Manage Multiple Ollama Models: 2026 Workflows]: /guides/managing-multiple-models-ollama/ - Manage multiple Ollama models in 2026: disk cleanup, switching, tagging. Qwen 3.6, Gemma 4, DeepSeek V4 (cloud-only) — practical workflows. [Beyond Transformers: 5 Architectures for Your $50 Mini PC]: /guides/beyond-transformers-5-architectures/ - We benchmarked RWKV-7 vs gemma3 on a $50 mini PC. The transformer crashed at turn 6. Here are 5 alternative architectures that run better on budget hardware. [Building a Distributed AI Swarm for Under $1,100]: /guides/build-distributed-ai-swarm-under-1100/ - A complete bill of materials for a three-node distributed AI cluster: RTX 3090 workstation, ThinkCentre M710Q for light inference, Raspberry Pi 5 coordinator. Every part sourced used or cheap, total cost under $1,100. [Building a Local AI Assistant: Your Private Jarvis]: /guides/building-local-ai-assistant/ - Build a private AI assistant with Ollama, Open WebUI, Whisper, and Kokoro TTS. Voice chat, document Q&A, home automation — all local, no cloud, no subscriptions. [Building AI Agents with Local LLMs: A Practical Guide]: /guides/local-ai-agents-guide/ - Build AI agents with local LLMs using Ollama and Python. Model requirements, VRAM budgets, framework comparison, working code example, and security warnings. [China Made Open Source a Strategy. If It Pulls Back, Who Fills the Gap?]: /guides/western-open-source-ai-china-gap/ - If China restricts future open weights, who fills the gap? The West's open-model capacity is real (Ai2, Mistral, Apertus) but scattered and underfunded. [China May Restrict Its AI Exports — Your Local Models Don't Care]: /guides/china-ai-export-controls-local-models/ - China's commerce ministry met Alibaba, ByteDance and Z.ai about walling off top AI models. The US did the same in June. The weights on your disk don't care. [Claude Code vs PI Agent — Which Coding Agent for Local AI?]: /guides/claude-code-vs-pi-agent-local-ai/ - Claude Code vs PI Agent compared for local AI development. System prompts, tools, pricing, local model support, and honest verdicts for every type of developer. [Claude Code's Source Just Leaked: What 500K Lines of TypeScript Reveal About AI Coding Agents]: /guides/claude-code-source-leak-what-we-learned/ - Claude Code's full source was exposed via npm source maps. Here's what the leaked architecture reveals about multi-agent orchestration, and what it means for local AI agent builders. [ClawHub Malware Alert: Top Skills Infected]: /guides/clawhub-malware-alert/ - The #1 ClawHub skill was malware. How it stole API keys, Cisco's new scanner tool, the MoldBot data leak, and 7 things to do right now to protect yourself. [CodeLlama vs DeepSeek Coder vs Qwen Coder: Best Local Coding Models Compared]: /guides/codellama-vs-deepseek-coder-vs-qwen-coder/ - CodeLlama vs DeepSeek Coder vs Qwen Coder vs Codestral benchmarked: HumanEval scores, VRAM per quant, and speed tests. Qwen 7B beats CodeLlama 70B. [ComfyUI Won — But A1111 Users Should Switch to Forge Neo Instead]: /guides/comfyui-vs-automatic1111-vs-fooocus/ - ComfyUI is faster, uses less VRAM, and gets new model support first. But the 2-3 week learning curve is real. If you're on A1111, Forge Neo gives you Flux support without starting over. Fooocus is dead. Speed tests and VRAM comparisons inside. [Context Length Exceeded: What To Do When Your Model Runs Out of Space]: /guides/context-length-exceeded-fix/ - Model forgetting earlier messages or throwing context errors? How context length works, what happens when it fills, and practical fixes for chat, RAG, and coding. [Context Length Explained: Why It Eats Your VRAM]: /guides/context-length-explained/ - What context length actually means for local LLMs, how it affects VRAM usage, practical limits for different hardware, and when you actually need 128K+ tokens. [ControlNet Guide: Precise AI Image Control on Your GPU]: /guides/controlnet-guide-beginners/ - ControlNet guide for Stable Diffusion and Flux. Covers Canny, OpenPose, Depth preprocessors, VRAM needs, ComfyUI and A1111 setup, and practical workflows. [CPU-Only LLMs 2026: Real tok/s, Best Models & a 70B Dual-Xeon Build]: /guides/cpu-only-llms-what-actually-works/ - No GPU? A decent CPU runs 7B models at 10-15 tok/s, and BitNet hits 45 tok/s in 0.4GB. Real benchmarks, best models, and a $1,100 dual-Xeon 70B build. [Crane + Qwen3-TTS: Run Voice Cloning Locally with Rust]: /guides/crane-qwen3-tts-local-voice-cloning/ - Clone any voice with 3 seconds of audio using Qwen3-TTS through Crane's pure Rust inference engine. ~4GB VRAM, faster than real-time, Apache 2.0. [CUDA Out of Memory: What It Means and How to Fix It]: /guides/cuda-out-of-memory-fix/ - CUDA out of memory means your model doesn't fit in VRAM. Seven fixes ranked by effort — context length, KV cache quantization, model quant, CPU offload — with tool-specific commands for Ollama, llama.cpp, and LM Studio. [DeepSeek Models Guide: R1, V3, and Coder]: /guides/deepseek-models-guide/ - Complete DeepSeek models guide covering R1, V3, and Coder locally. Which distilled R1 to pick for your GPU, VRAM requirements, and benchmarks vs Qwen 3. [DeepSeek V3.2 Guide: What Changed and How to Run It Locally]: /guides/deepseek-v3-2-guide/ - DeepSeek V3.2 was the Feb 2026 flagship — V4 now leads. But the R1-Distill models run on a $200 used GPU and remain the local reasoning pick. [DeepSeek V4 Flash vs Pro: Verdict, Cost, Setup (2026)]: /guides/deepseek-v4-flash-vs-pro-guide/ - Flash or Pro? Which DeepSeek V4 to run, what each costs (Haiku-tier pricing), and how to deploy locally. The verdict, the tradeoffs, no launch recap. [DeepSeek V4: Everything We Know Before It Drops]: /guides/deepseek-v4-preview/ - DeepSeek V4 launches next week with native image and video generation, 1M context, and rumored 1T MoE params with only 32B active. Here's what local AI builders need to know and how to prepare. [DFlash vs MTP on RTX 3090: I Tested Both Locally]: /guides/dflash-vs-mtp-rtx-3090-head-to-head/ - Firsthand head-to-head bench of DFlash + DDTree against MTP (PR #22673) on a single RTX 3090, same Qwen 3.6-27B target. Real numbers, both backends. [Distilled vs Frontier Models for Local AI — What You're Actually Getting]: /guides/distilled-vs-frontier-models-local-ai/ - That local model you love was probably trained on stolen outputs from Claude or GPT. Here's what distillation actually does to a model's reasoning, where it breaks, and why it matters most for agentic work. [Embedding Models for RAG: Which to Run Locally]: /guides/embedding-models-rag/ - nomic-embed-text is still the default for most local RAG setups — 274MB, 8K context, runs on CPU. But Qwen3-Embedding 0.6B just changed the game. Model picks, VRAM needs, speed numbers, and the chunking mistakes that break retrieval. [epsiclaw: OpenClaw Stripped to 515 Lines of Python (The Karpathy Treatment)]: /guides/epsiclaw-minimal-openclaw-515-lines/ - epsiclaw is a minimal, readable reimplementation of OpenClaw in 515 lines of Python with 6 files and one dependency. Inspired by Karpathy's approach to autoresearch. Here's what it does and why it matters. [Every SSD-Streaming MoE Engine: What's Real, What's Dead]: /guides/ssd-streaming-moe-engines/ - Eight engines that stream MoE experts from disk appeared in four months. Three have no license file at all. Here's the verified state of each. [Fine-Tuning LLMs on Consumer Hardware: LoRA and QLoRA Guide]: /guides/fine-tuning-local-lora-qlora/ - Fine-tune a 7B model on 6-10GB VRAM with QLoRA and Unsloth (2-5x faster, 70% less memory). Only 200-500 examples needed — and in 2026 you can fine-tune a 30B MoE on a single 24GB card. Dataset prep through training on RTX 3060-5090. [Fine-Tuning on Mac: LoRA & QLoRA with MLX]: /guides/fine-tuning-mac-lora-mlx/ - Fine-tune Llama, Qwen, and Mistral on Apple Silicon using mlx-lm. Real memory numbers, step-by-step commands, and how to deploy your model with Ollama. [Fix OpenClaw Token Waste: $150 to $6 Overnight]: /guides/openclaw-token-optimization/ - Cut OpenClaw API costs by 97% with three proven fixes: route heartbeats through Ollama, add tiered model routing, and purge session history token bloat. [Fix ROCm Not Detecting Your AMD GPU (2026)]: /guides/rocm-not-detecting-gpu-amd-fix/ - AMD GPU not detected in ROCm? Check supported GPUs, fix rocminfo errors, HSA_OVERRIDE hack for unsupported cards, and Ollama/llama.cpp ROCm build fixes. [Flash-MoE: What a 397B Model on a Laptop Actually Cost]: /guides/flash-moe-run-397b-model-laptop/ - Flash-MoE ran Qwen3.5-397B on a 48GB MacBook at 4.4 tok/s — then stopped dead in March 2026, unlicensed. The 2-bit trap it documented is the part worth keeping. [FP4 Just Landed in llama.cpp: NVFP4 vs MXFP4 Explained (2026)]: /guides/fp4-inference-llamacpp-nvfp4-mxfp4/ - NVFP4 in llama.cpp, MXFP4 in ik_llama.cpp. The first practical FP4 quantization for the GGUF ecosystem — what works, what doesn't, and what to test. [Free Local AI vs Paid Cloud APIs: Real Cost Comparison]: /guides/local-ai-vs-cloud-api-cost/ - A used RTX 3090 is $1,200 now and costs ~$10/month to run. Full break-even math vs OpenAI, Anthropic and Google APIs — including when local never pays back. [GB10 Boxes Compared: vs Strix Halo, vs Used 3090 (2026)]: /guides/gb10-boxes-compared/ - DGX Spark went $3,999 → $4,699 in Feb. The GB10 field, Strix Halo as a Windows alternative, and when a used 3090 beats all of them on speed. [Gemma 4 26B in 2GB RAM: The MoE Memory Ladder Explained]: /guides/gemma-4-2gb-ram-ssd-streaming/ - One model, three places its experts can live: VRAM, RAM, SSD. We measured the first two on Gemma 4. TurboFieldfare just added the third. [Gemma 4 Just Dropped: What Local AI Builders Need to Know]: /guides/gemma-4-local-ai-guide/ - Google's Gemma 4 is here -- dense and MoE variants, Apache 2.0, multimodal with vision and audio. VRAM requirements, benchmarks, and how it compares to Qwen 3.5. [Gemma Models Guide: Google's Lightweight Local LLMs]: /guides/gemma-models-guide/ - Gemma 3 27B beats Gemini 1.5 Pro on benchmarks and runs on a single GPU. The 4B outperforms Gemma 2 27B. Full lineup from 1B to 27B with VRAM needs, speeds, and honest comparisons. [GGUF File Won't Load: Format and Compatibility Fixes]: /guides/gguf-file-wont-load-fix/ - GGUF model won't load? Version mismatch, corrupted download, wrong format, split files, or memory issues. Find your error and fix it in under a minute. [GPT-OSS Guide: OpenAI's First Open Model for Local AI]: /guides/gpt-oss-guide-openai-local/ - GPT-OSS 20B is OpenAI's first open-weight model. MoE with 3.6B active params, MXFP4 at 13GB, 128K context, Apache 2.0. Here's how to run it. [GPU Buying Guide for Local AI: Pick the Right Card]: /guides/gpu-buying-guide-local-ai/ - The complete GPU buying guide for local AI. Covers RTX 3060 through 4090 with VRAM analysis, performance benchmarks, prices, and used vs new buying advice. [GTX 1650 vs RTX 3060 on a 35B MoE: What the Card Buys]: /guides/qwen3-6-35b-a3b-gtx-1650-4gb-vs-rtx-3060/ - A $60-class GTX 1650 4 GB runs Qwen3.6-35B-A3B at 20 tok/s on 32 GB of RAM. The RTX 3060 in the same slot does 28 at the same setting and 39 tuned. Measured. [Home Assistant Voice Control, No Cloud or Alexa (2026)]: /guides/home-assistant-local-llm-guide/ - Fully local voice control with Home Assistant, Ollama, Whisper, and Piper — no Alexa, no cloud, no subscriptions. Wyoming pipeline, model picks, hardware. [How Much Does It Cost to Run LLMs Locally?]: /guides/cost-to-run-llms-locally/ - $200-800 for hardware, $5-15/month in electricity, and a 3-6 month breakeven vs ChatGPT Plus at $240/year. Full cost breakdown with real numbers. [How OpenClaw Actually Works: Architecture Guide]: /guides/how-openclaw-works/ - 5 input types explain the 'alive' behavior: messages, heartbeats, crons, hooks, and webhooks feed a single agent loop. The 3am phone call was just a timer event. [How to Fix Slow Qwen 3.6 27B on RTX 3090 (10-80 tok/s)]: /guides/fix-slow-qwen-3-6-27b-rtx-3090/ - Qwen 3.6-27B at 12 tok/s on a 3090 when others report 35? The 8-step diagnostic checklist for offload, quants, templates, power limits, and backend choice. [How to Get 2.5x Faster Qwen on RTX 3090 (Free)]: /guides/dflash-rtx-3090-bench-both-qwens/ - I built DFlash on my RTX 3090 and ran the full bench. Real 2.5x speedup on Qwen 3.5 and 3.6 — below the 3.43x README claim, still huge. Here's how. [How to Run GLM 5.2 Locally: GPU, VRAM & Quant Guide]: /guides/run-glm-5-2-locally/ - GLM 5.2 is 753B params and 1.51TB at full precision. Run it locally: the live Unsloth quant ladder, every GPU and RAM path, and the quant to actually target. [How to Run Karpathy's Autoresearch on Your Local GPU]: /guides/karpathy-autoresearch-local-gpu-guide/ - Set up Karpathy's autoresearch on your GPU to run 100+ ML experiments overnight. Works on RTX 3090/4090 as-is, scales down to 6GB cards with tweaks. [How to Update Models in Ollama — Keep Your Local LLMs Current]: /guides/update-models-ollama/ - Ollama doesn't auto-update models. Run ollama pull model:tag to grab the latest version — only changed layers download. Use ollama show to check what you have, and a simple loop to update everything at once. [Hugging Face Hacked by AI Agent — Saved by a Local Model (2026)]: /guides/hugging-face-ai-agent-breach/ - Hugging Face disclosed July 16 that an autonomous AI agent breached its internal infra — now confirmed as OpenAI's own models. The models you download are safe — here's what was and wasn't hit. [I Got the AI Result I Wanted. Then I Ran It Nine More Times]: /guides/agent-skill-compilation-tested/ - A frontier model read my logs and wrote a skill that beat my local 27B baseline by 10.6 points. Nine more compilation runs showed the number was fake. [Inkling 975B vs Your 3090: The Real Memory Math (2026)]: /guides/open-frontier-vs-consumer-hardware/ - Inkling's smallest 1-bit quant is 270GB. A maxed consumer desktop holds 256GB. The open frontier left consumer hardware behind. Here's the honest math. [Intel Arc B580 for Local LLMs: 12GB VRAM at $250, With Caveats]: /guides/intel-arc-b580-local-llm/ - The Arc B580 gives you 12GB VRAM for $250, but Intel's AI software stack needs work. Real tok/s benchmarks, setup paths, and honest comparison with RTX 3060. [Intel Arc GPUs for Local AI: The Underdog Option That Actually Works]: /guides/intel-arc-local-ai/ - The Arc A770 16GB gives you 16GB of VRAM for ~$250 used. Software support through IPEX-LLM and llama.cpp SYCL is real but rough. Honest benchmarks, what works, and what doesn't. [Intel's $949 GPU Has 32GB VRAM and 608 GB/s Bandwidth: What It Means for Local AI]: /guides/intel-32gb-vram-gpu-local-ai/ - Intel is launching a 32GB VRAM GPU for $949. Here's how it compares to the RTX 3090, RTX 4090, and used GPU market for running local LLMs and Stable Diffusion. [Intent Engineering for Local AI Agents: A Practical Guide]: /guides/intent-engineering-local-ai-guide/ - Stop telling your agent to 'be helpful.' Start encoding specific goals, decision boundaries, and value hierarchies it can actually act on. Starter template included. [Is LM Studio Infected? How to Check Your Install (March 2026)]: /guides/lm-studio-malware-security-check/ - Reports of possible malware in LM Studio are circulating on Reddit. Here's what we know, how to verify your installation, and what to do if you're affected. [Is Qwen Going Closed? Open Weights vs Frontier (2026)]: /guides/qwen-open-weights-vs-closed-frontier-2026/ - Qwen split into a closed frontier (Max, Plus, VLA) and an open mid-tier (3.6-27B and 35B-A3B). The 3.7 open weights aren't here yet. The honest read. [Kimi K3 & Qwen 3.8: Open Weights You Can't Run (2026)]: /guides/open-weights-you-cant-run/ - Kimi K3's 2.8T weights need 64 accelerators to load. Qwen 3.8's 27B did ship, Apache 2.0, and fits 24GB. Openness and runnability are separate axes. [KV Cache: Why Context Length Eats Your VRAM (And How to Fix It)]: /guides/kv-cache-optimization-guide/ - The KV cache is why your 8B model OOMs at 32K context. Full formula, worked examples, TurboQuant, hybrid attention, and 7 ways to cut it. May 2026. [Laptop vs Desktop for Local AI: Which Should You Buy?]: /guides/laptop-vs-desktop-local-ai/ - A $1,200 desktop RTX 3090 gives you 24GB VRAM. The same money in a gaming laptop gets 8GB. MacBooks break the rules with 48GB+ unified memory for 70B models. [LightClaw: A 7,000-Line Python Alternative to OpenClaw]: /guides/lightclaw-lightweight-openclaw-alternative/ - OpenClaw is 40,000+ lines of TypeScript. LightClaw does Telegram AI assistant, 6 LLM providers, memory, skills, and agent delegation in ~7,000 lines of Python. One week old, 12 stars, one developer. Here's what it can and can't do. [Lightning 2.6.x Malware: Check Your Local AI Stack]: /guides/pytorch-lightning-malware-local-ai-audit/ - PyPI's lightning package was poisoned April 30 with malware that abuses Claude Code hooks. Here's the 5-minute audit I ran on my own 3090 box. [LiquidAI LFM2: The First Hybrid Model Built for Your Hardware]: /guides/liquidai-lfm2-local-setup-guide/ - LFM2-24B-A2B runs at 112 tok/s on CPU with only 2.3B active params. Not a transformer. GGUF files from 13.5GB, Ollama and llama.cpp setup, and where it beats Qwen. [Llama 3 Guide: Every Size from 1B to 405B]: /guides/llama-3-guide-every-size/ - Complete Llama 3 guide covering every model from 1B to 405B. VRAM requirements, Ollama setup, benchmarks vs Qwen 3, and which size fits your hardware. [Llama 4 Guide: Running Scout and Maverick Locally (2026)]: /guides/llama-4-guide-scout-maverick/ - Complete Llama 4 Scout (109B MoE) and Maverick guide for local AI. VRAM, Ollama and vLLM setup, hardware reality, and how it stacks against Qwen 3.6. [Llama 4 vs Qwen3 vs DeepSeek V3.2: Which to Run Locally in 2026]: /guides/llama-4-vs-qwen3-vs-deepseek-v3-2-local/ - Llama 4 needs 55GB. DeepSeek V3.2 needs 350GB. Qwen3 runs on 8GB. Here's who wins at each VRAM tier and use case for local AI in 2026. [llama.cpp Build Errors: Common Fixes for Every Platform]: /guides/llamacpp-build-errors-fixes/ - llama.cpp won't build or runs wrong? CMake, CUDA, Gemma 4 thinking-mode, Qwen 3.6 kwargs, num_ctx VRAM overflow. Exact fixes for every platform. [llama.cpp Just Got a New Home: What the Hugging Face Acquisition Means for Local AI]: /guides/llamacpp-hugging-face-ggml-acquisition/ - ggml.ai — the team behind llama.cpp — is joining Hugging Face. Open source stays open, Georgi keeps the wheel. What changed, what didn't, and what to watch. [llama.cpp vs Ollama vs vLLM: One User vs Many (2026)]: /guides/llamacpp-vs-ollama-vs-vllm/ - Single-user, the three are closer than benchmark posts admit. Concurrent, vLLM pulls 10-20x ahead. Decision tree, the vLLM VRAM gotcha, mid-2026 versions. [LLM Running Slow? Two Different Problems, Two Different Fixes]: /guides/llm-running-slow-fix/ - Slow local LLM? Separate time-to-first-token from generation speed. Fix prompt processing with batch size and Flash Attention. Fix tok/s with GPU layers, quantization, and context length. [LM Studio Tips & Tricks: Faster + Features You Miss (2026)]: /guides/lm-studio-tips-and-tricks/ - Run local LLMs faster in LM Studio with speculative decoding and MLX, plus the API server, GPU offload, and power-user features most people never touch. [LM Studio vs llama.cpp: Why Your Model Runs Slower in the GUI]: /guides/lm-studio-vs-llamacpp-speed-gap/ - LM Studio uses llama.cpp under the hood but often runs 30-50% slower. Bundled runtime lag, UI overhead, and default settings explain the gap. How to benchmark it yourself and when the convenience is worth it. [LM Studio vs Ollama on Mac: Which Should You Use?]: /guides/lm-studio-vs-ollama-mac/ - LM Studio's MLX backend is 20-30% faster and uses half the memory. Ollama is lighter, always-on, and better for APIs. Mac-specific benchmarks and when to use each. [Local AI for Accounting and Tax: Keep Your Financial Data Off the Cloud]: /guides/local-ai-accounting-tax-privacy/ - Local LLMs can categorize transactions, draft client letters, extract receipt data, and answer questions over tax documents — without sending a single number to OpenAI or Google. What works, what doesn't, and how to set it up. [Local AI for Lawyers: Confidential Document Analysis Without Cloud Risk]: /guides/local-ai-for-lawyers/ - A federal judge ordered OpenAI to hand over 20 million chat logs. If you're a lawyer using ChatGPT for client work, that's an ethics problem. Local AI keeps everything on your hardware. [Local AI for Privacy: What's Actually Private]: /guides/local-ai-privacy-guide/ - Running AI locally keeps prompts off corporate servers — but model downloads, telemetry, and VS Code extensions can still leak data. Here's what's genuinely private, what isn't, and how to close every gap. [Local AI for Small Business: Email, Invoicing, and Customer Support Without Monthly Subscriptions]: /guides/local-ai-small-business-replace-subscriptions/ - A 5-person team spends $1,500-3,000/year on AI subscriptions. A $600 mini PC running Ollama replaces all of them. Here's the setup, the workflows, and the math. [Local AI for Therapists: Session Notes, Treatment Plans, and Client Privacy Without the Cloud]: /guides/local-ai-for-therapists/ - Run AI on your own hardware to draft session notes, treatment plans, and clinical letters without sending client data to OpenAI. HIPAA-friendly setup for therapists. [Local AI Troubleshooting Guide: Every Common Problem and Fix]: /guides/local-ai-troubleshooting-guide/ - Model running 30x slower than expected? Probably on CPU instead of GPU. Fixes for won't-load errors, CUDA crashes, garbled output, and OOM across Ollama and LM Studio. [Local AI Upscaling: Make Blurry Images Sharp Without the Cloud]: /guides/local-ai-upscaling-guide/ - Upscayl, Real-ESRGAN, chaiNNer, and ComfyUI can upscale your photos for free on your own hardware. No subscriptions, no uploads, no per-image fees. Even a GTX 1060 works. Here's how to pick the right tool and start. [Local LLMs vs ChatGPT: An Honest Comparison]: /guides/local-llms-vs-chatgpt-honest-comparison/ - ChatGPT has web search, voice mode, and GPT-5.2. Local LLMs have privacy, no subscriptions, and no rate limits. Here's when each one wins, what the cost math actually looks like, and why most power users run both. [Local LLMs vs Claude: When Each Actually Wins]: /guides/local-llms-vs-claude/ - Qwen 3 32B matches Claude on daily tasks at zero marginal cost. Claude still wins on 200K-token documents and multi-step debugging. Benchmarks, pricing, and when to use each. [Local RAG: Search Your Documents with a Private AI]: /guides/local-rag-search-documents-private-ai/ - Search your private PDFs, notes, and code with a local LLM—no cloud, no API calls. 3 setup methods from zero-config Open WebUI to 30 lines of Python with ChromaDB. [LocalAgent: A Local-First Agent Runtime That Actually Cares About Safety]: /guides/localagent-local-first-agent-runtime-safe-tool-calling/ - Rust CLI for AI agents with deny-by-default permissions, approval workflows, and deterministic replay. Works with LM Studio, Ollama, and llama.cpp. [LoRA Skill Compilation Is a Double-Headed Coin Flip]: /guides/skill-compilation-into-weights-ten-seeds/ - Ten LoRA seeds on identical data spread 3.62 points, against 4.65 for prompt compilation. Not one beat the no-adapter base. Measured on an RTX 3090. [LoRA Training on Consumer Hardware: Fine-Tune Models With 12GB VRAM]: /guides/lora-training-consumer-hardware/ - QLoRA fine-tunes a 7B model on an RTX 3060 12GB in 2-4 hours. Full Unsloth and Axolotl recipes, VRAM tables, and the GGUF export pipeline. [M4 Max and M3 Ultra for Local LLMs: Apple Silicon in 2026]: /guides/m4-max-ultra-local-llms-apple-silicon/ - No M4 Ultra exists. After Apple's 2026 memory cuts, the Mac Studio pairs the M4 Max (64GB) with the M3 Ultra (96GB, 819 GB/s). Which to buy for local AI. [Mac Mini M4 for Local AI: Which Config to Buy and What It Actually Runs]: /guides/mac-mini-m4-local-ai/ - Mac Mini M4 Pro 48GB runs Qwen 3.6-35B-A3B silently at 40W. Which config to buy after Apple's 2026 price hikes, and what each tier actually runs for local AI. [Mac Studio for Local AI: Is It Worth the Price?]: /guides/mac-studio-local-ai-workstation/ - Mac Studio M4 Max (64GB) and M3 Ultra (96GB) for local LLMs after Apple's 2026 memory cuts. Real tok/s, cost vs dual RTX 3090, and who should buy one. [Mac vs PC for Local AI: Which Should You Choose?]: /guides/mac-vs-pc-local-ai/ - An RTX 3090 runs 7B-32B models 2-3x faster than a Mac. A 96GB Mac Studio or a Strix Halo mini-PC (from ~$1,499) loads 70B. Benchmarks, current 2026 prices, and which platform fits. [Memory Leak in Long Conversations: Causes and Fixes]: /guides/memory-leak-long-conversations-fix/ - VRAM climbs with every message until your model crashes? It's probably KV cache growth, not a leak. How to diagnose, monitor, and fix memory issues in local LLMs. [Mistral Voxtral TTS: Open-Weight Voice AI You Can Run Locally]: /guides/mistral-voxtral-tts-local-voice-ai/ - Voxtral TTS is a 4B open-weight text-to-speech model that beats ElevenLabs Flash v2.5 in blind tests. 70ms latency, 9 languages, voice cloning from 3 seconds. Here's how to run it. [Mixtral 8x7B & 8x22B VRAM Requirements]: /guides/mixtral-8x7b-8x22b-vram-requirements/ - Mixtral 8x7B and 8x22B VRAM requirements at every quantization level — plus why the 'every expert must sit in VRAM' rule Mixtral taught no longer holds in the A3B era. [Mixtral VRAM Requirements: 8x7B and 8x22B at Every Quantization Level]: /guides/mixtral-vram-requirements/ - Mixtral 8x7B has 46.7B params but only 12.9B activate per token. You still need VRAM for all 46.7B. Exact VRAM for every quant from Q2 to FP16. [Model Formats Explained: GGUF vs GPTQ vs AWQ vs EXL2]: /guides/model-formats-explained-gguf-gptq-awq-exl2/ - GGUF vs GPTQ vs AWQ vs EXL2/EXL3 model formats explained — plus the new FP4 (MXFP4/NVFP4) and Apple's MLX. What each does, which tools run it, and how to choose for your GPU. [Model Outputs Garbage: Debugging Bad Generations]: /guides/model-outputs-garbage-debug/ - Local LLM outputs repetitive loops, gibberish, or wrong answers? Seven causes with exact fixes — from corrupted downloads to wrong chat templates. [Model Routing for Local AI — Stop Using One Model for Everything]: /guides/model-routing-local-ai-guide/ - You're running one model for every task. That wastes VRAM, burns electricity, and gives worse results. Model routing sends each task to the right model at the right cost. Here's how to set it up. [MoE Models Explained: Why Mixtral Uses 46B Parameters But Runs Like 13B]: /guides/moe-models-explained/ - MoE explained with our own 3090 and 3060 numbers: a 35B MoE fits 24 GB and beats dense by up to 4x, offload moved the wall to RAM, and where dense still wins. [Multi-GPU Setups for Local AI: Worth It?]: /guides/multi-gpu-worth-it/ - Dual RTX 3090s cost $1,700-2,000 and need a 1,200W PSU — but a single 3090 at ~$1,000 runs every model under 32B. When two GPUs actually beat one bigger card, and when they don't. [mycoSwarm vs Exo vs Petals vs Nanobot: What's Actually Different]: /guides/mycoswarm-vs-exo-petals-nanobot/ - Exo distributes inference across Macs. Petals shares GPUs with strangers. Nanobot routes your queries to Chinese clouds without asking. The real question: who controls where your prompts go? [Mythos AI Cracked Apple's Best Defense in 5 Days]: /guides/mythos-cracked-apple-m5-5-days/ - Anthropic's Mythos helped Calif bypass Apple's Memory Integrity Enforcement on M5 in 5 days — what Apple spent 5 years building. The offense-defense compression is real. [nanollama: Train Your Own Llama 3 From Scratch on Custom Data]: /guides/nanollama-train-llama-from-scratch/ - Pretrain Llama 3 architecture models from raw text, export to GGUF, and run with llama.cpp. Forked from Karpathy's nanochat. 46M to 7B parameters. [NVIDIA GPU Prices Are Rising: What to Do Now]: /guides/nvidia-gpu-prices-rising/ - GPU prices are spiking due to GDDR7 shortages and AI datacenter demand. Here's what's happening, which cards are affected, and strategies for local AI builders. [Nvidia PAIR Bought Me 7 Percent. My Own Router Cost Me Half.]: /guides/nvidia-pair-measured-rtx-3090-3060/ - Nvidia's new home-network AI router split 20 requests across an RTX 3090 and a 3060 for 1.07x over the 3090, then dropped half a 27B queue. Mine: 0.52x. [Obsidian + Local LLM: Build a Private AI Second Brain]: /guides/obsidian-local-llm-guide/ - Connect Obsidian to a local LLM via Ollama for private AI-powered note search, summaries, and chat. Step-by-step setup with Copilot and Smart Connections. [Ollama 0.30.0: What's New, What's Faster, What Breaks on Upgrade]: /guides/ollama-0-30-0-whats-new/ - Ollama 0.30.0: llama.cpp integration, flash-attention default for Qwen/Gemma, broader model support. Firsthand upgrade notes, known issues to watch. [Ollama API Connection Refused: Quick Fixes]: /guides/ollama-api-connection-refused-fix/ - Ollama API returning connection refused? Check if it's running, fix the port, open it to the network, and solve Docker and WSL2 connectivity issues. [Ollama Not Using GPU: Complete Fix Guide (2026)]: /guides/ollama-not-using-gpu-fix/ - Ollama running on CPU instead of GPU? Diagnose with ollama ps and nvidia-smi, then fix CUDA, ROCm, WSL2, VRAM limits, and the new Ollama Cloud confusion. [Ollama on Mac Not Working? Fix Metal, Memory Pressure, and Slow Performance]: /guides/ollama-mac-troubleshooting/ - ollama ps says CPU? Generation crawling at 2 tok/s? macOS killed your model mid-sentence? Every Mac-specific Ollama problem diagnosed and fixed with exact commands. [Ollama on Mac: Setup and Optimization Guide (2026)]: /guides/ollama-mac-setup-optimization/ - Install Ollama on Apple Silicon, verify Metal GPU is active, and tune it for your Mac's RAM. Config for M1 through M4 Ultra with model picks per memory tier. [Ollama Troubleshooting Guide: Every Common Problem and Fix]: /guides/ollama-troubleshooting-guide/ - GPU not detected? Running at 1/30th speed on CPU? OOM crashes mid-generation? Every common Ollama error with exact diagnostic commands and fixes for Mac, Windows, and Linux. Updated July 2026 for v0.31.x and Qwen 3.5 + 3.6. [Ollama vs LM Studio: Speed, Setup, and Verdict]: /guides/ollama-vs-lm-studio/ - Ollama gives you a CLI with 100+ models and an OpenAI-compatible API. LM Studio gives you a visual GUI with one-click downloads. Most power users run both—here's when to use each. [Open WebUI Not Connecting to Ollama? Every Fix]: /guides/open-webui-ollama-connection-fix/ - Docker networking, wrong OLLAMA_BASE_URL, localhost confusion, WSL2 isolation, missing models, random disconnects. Every Open WebUI + Ollama connection problem with the exact fix. [Open WebUI Setup Guide: ChatGPT UI for Local AI]: /guides/open-webui-setup-guide/ - 1 Docker command gives you a ChatGPT-like interface for any Ollama model. 120K+ GitHub stars, built-in RAG, voice chat, and multi-model switching—all running locally. [OpenClaw After Steinberger — What the OpenAI Move Means for Your Setup]: /guides/openclaw-after-steinberger-what-changes/ - Peter Steinberger joined OpenAI. Three releases shipped since. Elon Musk posted a meme. Baby Keem is debugging agents. Here's what actually matters for your OpenClaw setup. [OpenClaw ClawHub Alert: 1,103 Malicious Skills Found]: /guides/openclaw-clawhub-security-alert/ - OpenClaw ClawHub security alert: 1,103 malicious skills found across 14,706 audited. CVE-2026-28458 Browser Relay auth bypass. How to protect yourself now. [OpenClaw Critical Sandbox Escape: Update to 2026.3.28 Now]: /guides/openclaw-security-report-march-2026/ - Ant AI Security Lab found 33 vulnerabilities in OpenClaw including critical privilege escalation and filesystem sandbox escape. If you're self-hosting, update immediately. [OpenClaw Memory Problems: Context Rot and the Forgetting Fix]: /guides/openclaw-memory-context-rot/ - Your OpenClaw agent forgets instructions, repeats questions, contradicts itself. How memory and context rot work — and fixes that apply to any harness. [OpenClaw Model Combinations: What to Pair for Each Task]: /guides/openclaw-best-model-combinations/ - Stop running one model for everything in OpenClaw. Pair Qwen 2.5 Coder 32B for autocomplete, Qwen 3.5 27B for planning, and Qwen3-Coder-Next for agentic coding. Combos by VRAM tier. [OpenClaw Model Routing: Cheap Models for Simple Tasks, Smart Models When Needed]: /guides/openclaw-model-routing/ - Stop paying Opus prices for heartbeats. Set up tiered model routing in openclaw.json so cheap models handle 80% of work and frontier models only fire when needed. [OpenClaw on Mac: Setup, Optimization, and What Actually Works]: /guides/openclaw-mac-setup-guide/ - brew install openclaw-cli, connect Ollama, configure the gateway, and stop fighting macOS. Apple Silicon setup, memory math, launchd config, and the gotchas nobody warns you about. [OpenClaw on Raspberry Pi: What Actually Works (and What Doesn't)]: /guides/openclaw-raspberry-pi/ - Pi 5 with 8GB RAM runs OpenClaw as a gateway with cloud APIs. Local LLMs hit 2-7 tok/s on 1.5B-3B models. Step-by-step setup for llama.cpp, Ollama, and OpenClaw on ARM64. [OpenClaw Security Guide: Risks and Hardening]: /guides/openclaw-security-guide/ - 42,000+ exposed instances, Google suspending accounts that connected via OAuth, 26% of ClawHub skills with vulnerabilities. Real risks, prompt injection demos, and step-by-step hardening for OpenClaw. [OpenClaw Security Hardening — Every Fix in February 2026]: /guides/openclaw-security-february-2026/ - SSRF bypass, sandbox escapes, unauthorized disk writes, session hijacking. Every security fix OpenClaw shipped in February 2026, explained in plain English. [OpenClaw Security Report: February 2026 — ClawHub Malware, Google Suspensions, and Critical Fixes]: /guides/openclaw-security-report-february-2026/ - 17 security fixes, 341 malicious ClawHub skills, Google banning users, and the creator leaving. Every OpenClaw security event from February 2026. [OpenClaw Security Report: January 2026]: /guides/openclaw-security-report-january-2026/ - Three high-severity CVEs, a supply chain attack on ClawHub, and 21,000+ exposed instances. Every OpenClaw security event from January 2026 with sources. [OpenClaw Trading Scams: How to Spot AI Agent Grifts Before They Cost You]: /guides/openclaw-trading-scams/ - AI agent trading scams target technical users who know agent pipelines are buildable. Here's the playbook they use, the math that breaks their claims, and how to protect yourself. [OpenClaw vs Commercial AI Agents: Which Should You Use?]: /guides/openclaw-vs-commercial-ai-agents/ - OpenClaw costs $0 plus API fees. Lindy runs $49-299/month but has 7,000+ integrations and SOC 2 compliance. Privacy, reliability, and customization compared honestly. [OpenClaw vs Cursor: Local AI Agent or Cloud IDE?]: /guides/openclaw-vs-cursor/ - OpenClaw is free, private, and runs your own models. Cursor is polished, fast, and cloud-powered. A developer's comparison: cost, privacy, model flexibility, offline use, and where each one wins. [OpenClaw's Creator Just Joined OpenAI — Here's What It Means for Local AI Agents]: /guides/openclaw-openai-acquihire-what-it-means/ - Peter Steinberger built the fastest-growing open-source project ever. Now he's at OpenAI. OpenClaw stays open. Here's what changes for local AI builders. [Ornith 1.5 35B vs Qwen 3.6 on RTX 3090: Speed Tested]: /guides/ornith-1-5-35b-vs-qwen-3-6-35b-rtx-3090/ - Firsthand A-B-B-A bench of Ornith 1.5-35B-A3B against Qwen 3.6-35B-A3B on one RTX 3090. Generation, prefill, VRAM, and the noise floor under all three. [Ouro-2.6B-Thinking: ByteDance's Looped Model That Punches Like an 8B]: /guides/ouro-2b-thinking-looped-language-model-local/ - Ouro-2.6B loops through the same transformer blocks 4 times to match 8B models at 2.6B parameters. Under 2GB at Q4. How the architecture works and why it matters. [PaddleOCR-VL: A 0.9B OCR Model That Runs on Any Potato]: /guides/paddleocr-vl-local-document-ocr/ - PaddleOCR-VL does document OCR — text, tables, formulas, charts — in 0.9B parameters. 109 languages. Now runs via llama.cpp and Ollama. Private, local, nearly free. [Phi Models Guide: Microsoft's Small but Mighty LLMs]: /guides/phi-models-guide/ - Phi-4 14B scores 84.8% on MMLU — matching models 5x its size — and fits on a 12GB GPU at Q4. The full Phi-4 lineup, including the reasoning and vision-reasoning variants, with VRAM needs, benchmarks, and honest weaknesses. [Pi AI vs Local AI: Cloud Companion or Private Assistant?]: /guides/pi-ai-vs-local-ai/ - Pi.ai is warm, free, and cloud-only. Local AI is private, flexible, and yours. What Pi does well, where it falls short, and when running your own model is the better call. [Quantization Explained: What It Means for Local AI]: /guides/llm-quantization-explained/ - Q4_K_M shrinks a 7B model from 14GB to ~4GB while keeping 90-95% quality. What every quantization format means, how much VRAM each saves, and which to pick for your GPU. [Qwen 3.5 Locally — 27B vs 35B-A3B vs 122B, Which Model Fits Your GPU]: /guides/qwen35-local-guide-which-model-fits-your-gpu/ - Qwen 3.5 and 3.6 on local hardware. 27B dense vs 35B-A3B MoE vs 122B compared. VRAM tables, community tok/s on RTX 3090, and which to pick for your card. [Qwen 3.5 Small Models: The 9B Beats Last-Gen 30B — Here's What Matters for Local AI]: /guides/qwen-3-5-small-models-9b-beats-30b/ - Alibaba's Qwen 3.5 drops 4 small models (0.8B to 9B) — all natively multimodal, 262K context, Apache 2.0. The 9B beats Qwen3-30B on reasoning and destroys GPT-5-Nano on vision. VRAM tables and what to run. [Qwen 3.6 Complete Guide: 27B Dense, 35B-A3B MoE, and Which to Use]: /guides/qwen-3-6-local-ai-guide/ - Qwen 3.6 landed in two open-weight flavors: 27B dense and 35B-A3B MoE. Benchmarks, hardware fit, and which variant to run on your GPU. [Qwen 3.6 MoE Routing, Measured: Flat Is the Wrong Number]: /guides/qwen-3-6-moe-expert-routing-measured/ - I traced every expert routing decision Qwen 3.6-35B-A3B makes across six workloads on an RTX 3060. Routing isn't flat, and 112 slots is the whole answer. [Qwen 3.6: Why Q4 Quant Breaks Local Coding Agents (And the Fix)]: /guides/qwen-3-6-q4-quant-coding-agents/ - A viral thread says Q4-to-Q6 fixes Qwen 3.6 coding, but the test was confounded. What four independent reports show about the quant tax on coding agents. [Qwen 3.7 Open Weights Watch: The June Window Is Closing]: /guides/qwen-3-7-preview-scored-57-aai-27b-35b-open-weights-watch/ - June 19: Qwen 3.7 open weights overdue. 3.5→3.6 cadence projected June 6-14; we're past. Max May 20 (56.6 AAI), VLA, Plus closed. Bench at drop. [Qwen 3.8 27B vs 3.6 on RTX 3090: Speed and Quality Tested]: /guides/qwen-3-8-27b-vs-3-6-27b-rtx-3090/ - Firsthand benchmarks of Qwen 3.8-27B against 3.6-27B on one RTX 3090: generation within a percent, VRAM +254 MiB, and HumanEval pass@1 a statistical tie. [Qwen vs Llama vs Mistral: Which Model Family Should You Build On?]: /guides/qwen-vs-llama-vs-mistral-model-shootout/ - Qwen has 201 languages and a model for every task. Llama has the biggest community. Mistral pioneered efficient MoE. Decision framework for choosing your model family in 2026. [Qwen2.5-VL Not Loading in LM Studio? Fix mmproj and Vision Errors]: /guides/qwen25-vl-lm-studio-troubleshooting/ - Fix every Qwen2.5-VL error in LM Studio: missing mmproj, 'model type not supported', no eye icon, vision crashes. Exact fixes with file paths. [Qwen3 Complete Guide: Every Model from 0.6B to 235B]: /guides/qwen3-complete-guide/ - Qwen3 is the best open model family for budget local AI. Dense models from 0.6B to 32B, MoE models that punch above their weight, and a /think toggle no one else has. [RAG Pipeline for Local AI: A Practical Guide to Retrieval-Augmented Generation]: /guides/rag-pipeline-local-ai-guide/ - Build a local RAG pipeline with Ollama, ChromaDB, and your own documents. Chunking strategies, embedding models, vector stores, and the failure modes nobody warns you about. [Razer AIKit Guide: Multi-GPU Local AI on Your Desktop]: /guides/razer-aikit-guide/ - Open-source Docker stack bundling vLLM, Ray, LlamaFactory, and Grafana into 1 container. Auto-detects GPUs, supports 280K+ HuggingFace models, and handles multi-GPU parallelism. [Replace GitHub Copilot With Local LLMs in VS Code — Free, Private, No Subscription]: /guides/replace-github-copilot-local-llms-vscode/ - Set up free, private AI code completion in VS Code with Continue + Ollama. Autocomplete, chat, and agentic coding with Qwen models at every VRAM tier. Step-by-step setup, model picks, honest tradeoffs. [ROCm vs CUDA for Local AI in 2026: The Software Gap Nobody Talks About]: /guides/rocm-vs-cuda-local-ai-2026/ - AMD GPUs have the bandwidth. They have the VRAM. They still lose by 2x on inference speed. Here's why, what actually works on ROCm 7.2, and whether RDNA 4 fixes anything. [RTX 3060 vs 3060 Ti vs 3070 for Local AI]: /guides/rtx-3060-vs-3060ti-vs-3070-local-ai/ - The RTX 3060 12GB now costs $90 MORE than the 3060 Ti and still wins for LLMs. Why the used market flipped, what $/GB reveals, and when the 3070 makes sense. [RTX 3090 vs 4070 Ti Super for Local LLMs]: /guides/rtx-3090-vs-4070-ti-super-local-llms/ - Head-to-head comparison of the RTX 3090 and RTX 4070 Ti Super for running LLMs locally. Covers VRAM, speed, power, price, and which to buy for your use case. [RTX 4090 vs Used RTX 3090 for Local AI: Which to Buy in 2026]: /guides/rtx-4090-vs-used-rtx-3090-local-ai/ - Both have 24GB VRAM. One costs about twice as much. RTX 4090 vs used RTX 3090 — real benchmarks, real prices, and who should buy which for local AI. [RTX 5060 Ti 16GB Killed? Local AI Alternatives]: /guides/rtx-5060-ti-16gb-local-ai-options/ - The RTX 5060 Ti 16GB faces production cuts from GDDR7 shortages. See what is really happening and explore the best alternative GPUs for local AI in 2026. [RTX 5060 Ti Review for Local AI — The New Budget King]: /guides/rtx-5060-ti-local-ai-benchmarks/ - Real benchmarks for the RTX 5060 Ti 16GB running local LLMs. Qwen 3.5 35B at 44 tok/s, 100K context for ~$430. Compared against RTX 3060, 3090, and 4060 Ti. [RTX 5090 Benchmarks: 5090 vs 4090 vs Used 3090 (2026)]: /guides/rtx-5090-local-ai-benchmarks/ - 5090 community benches across 4K-131K context, prompt-processing tables, 5090-vs-4090 upgrade math, and InsiderLLM's firsthand 3090 honest-value anchor. [RTX 5090 for Local AI: Worth the Upgrade?]: /guides/rtx-5090-local-ai-worth-it/ - 32GB GDDR7, 1,792 GB/s bandwidth, 67% faster than 4090 — but ~$3,900-4,400 street (Aug 2026). Full benchmarks, value analysis, and who should actually buy one. [Run AI Offline: Air-Gapped LLMs, No Internet (2026)]: /guides/running-ai-offline-complete-guide/ - Ollama runs fully offline after one download — disconnect the network and your AI keeps working. No accounts, no APIs. Setup, offline RAG, portable kits. [Run LLMs on Mac M-Series: Faster, Without the Gotchas (2026)]: /guides/running-llms-mac-m-series/ - Foundational how-to for Apple Silicon local AI: unified memory, MLX vs Ollama vs llama.cpp Metal, verification, and the headless Mac Mini AI server. [Run LLMs on Old Phones: A Practical Guide to Mobile AI Inference]: /guides/run-llms-old-phones-mobile-inference/ - That old Pixel 6 or Galaxy S21 in your drawer can run a local LLM. Realistic tok/s by phone tier, Termux setup, app options, and an honest phone vs Raspberry Pi comparison. [Run Qwen2.5-VL Vision in LM Studio (Setup)]: /guides/qwen25-vl-lm-studio-vision-setup/ - Get Qwen2.5-VL running in LM Studio in 5 minutes. Covers the mmproj file most people miss, correct download links, and how to analyze images and PDFs locally. [Run Your First Local LLM in 15 Minutes]: /guides/run-first-local-llm/ - Install Ollama, pull a model, and chat with AI offline—all in 15 minutes. Works on any Mac, Windows, or Linux machine with 8GB RAM. No accounts, no API keys, no fees. [Running 70B Models Locally — Exact VRAM by Quantization]: /guides/running-70b-models-locally-vram-guide/ - Llama 3.3 70B needs 43GB at Q4, 75GB at Q8, 141GB at FP16. Every quant level, which GPUs fit, real speeds, and when a 27B or MoE model is the smarter buy. [Running OpenClaw 100% Local — Zero API Costs]: /guides/openclaw-local-zero-api-costs/ - Configure OpenClaw to run entirely through Ollama with no API keys, no cloud calls, and no monthly bills. Full setup guide with model picks by VRAM tier. [Running OpenClaw on 4GB, 6GB, and 8GB GPUs: What Actually Works]: /guides/openclaw-low-vram-gpus/ - OpenClaw on low VRAM GPUs: 4GB is rough, 6GB is marginal, 8GB is where it starts working. Model picks, quantization tricks, partial offload, and when to just use a cloud API instead. [RWKV-7: Infinite Context, Zero KV Cache — The Local-First Architecture]: /guides/rwkv-7-local-ai-guide/ - RWKV-7 uses O(1) memory per token. Context length doesn't increase VRAM. At all. 16 tok/s on a Raspberry Pi. Here's why it matters for local AI and how to run it. [SDXL vs SD 1.5 vs Flux: Which Image Model Should You Run Locally?]: /guides/sdxl-vs-sd-1-5-vs-flux/ - SDXL vs SD 1.5 vs Flux compared by VRAM, speed, and quality. SD 1.5 needs 4GB, SDXL needs 8GB, Flux needs 12GB+. Benchmarks on real GPUs inside. [Session-as-RAG: Teaching Your Local AI to Actually Remember]: /guides/session-as-rag-local-ai-memory/ - Build persistent conversation memory for local LLMs. Chunk sessions, embed in ChromaDB, retrieve relevant past exchanges at query time. Full Python implementation with topic splitting and date citations. [Skills in the Weights: The LoRA Answer to the Prompt Tax]: /guides/skills-in-weights-lora-prompt-tax/ - Compiling a skill into a prompt costs 1,383.9 tokens per call, forever. Putting it in a LoRA costs 7.55 GB and a training run. Only one has been measured. [Slash Your AI Costs With a Token Audit]: /guides/token-audit-guide/ - Your AI API bill is higher than it needs to be. A 15-minute token audit finds the waste — system prompts, ballooning history, hidden tool tokens. Here's the exact process. [SmarterRouter: A VRAM-Aware LLM Gateway for Your Local AI Lab]: /guides/smarterrouter-vram-aware-llm-gateway-local-ai/ - Intelligent router that profiles your models, manages VRAM, caches responses semantically, and auto-picks the best model per prompt. Works with Ollama and llama.cpp. [Speculative Decoding: Free 20-50% Speed Boost for Local LLMs]: /guides/speculative-decoding-explained/ - Speculative decoding uses a small draft model to predict tokens verified by the big model. Same output, 20-50% faster. Setup guide for LM Studio and llama.cpp. [Stable Diffusion Locally: Getting Started]: /guides/stable-diffusion-locally-getting-started/ - SD 1.5 runs on 4GB VRAM, SDXL needs 8GB, Flux needs 12GB+. Generate unlimited images for free in under 5 minutes with Fooocus or ComfyUI. Setup, models, and first image tips. [Stable Diffusion on Mac: Image Generation with MLX and Draw Things]: /guides/stable-diffusion-mac-mlx/ - Draw Things generates SD 1.5 images in 8-15 seconds on an M2 Pro. ComfyUI takes 3x longer. MLX is fastest but code-only. Complete Mac image gen guide with speed tests. [Talk to Your Local LLM: Voice Chat Setup]: /guides/voice-chat-local-llms-whisper-tts/ - Under 1 second response time with Whisper + Kokoro TTS + your local model. Full setup guide for Open WebUI voice chat and standalone options. Needs 2-4GB VRAM. [Tesla P40: Still the Cheapest 24GB Card for Local AI]: /guides/used-tesla-p40-local-ai/ - 24GB VRAM for $220-250 bare on eBay (Sep 2026). Pascal architecture, no display output, passive cooling. Full benchmarks, setup guide, and honest comparison to the RTX 3060 and 3090. [Text Generation WebUI Setup Guide (2026)]: /guides/text-generation-webui-oobabooga-guide/ - Install oobabooga's TextGen (formerly text-generation-webui), load GGUF/EXL2/EXL3 models, and configure GPU offloading. Now with vision, tool-calling, and an Anthropic-compatible API. Covers the settings most guides skip. [The $36 RAM Fix That Made CPU Inference 56% Faster]: /guides/dual-channel-ram-local-llm-cpu-inference/ - Adding a second RAM stick to a mini PC lifted CPU token generation 52-58% across four models. Prompt processing moved under 2%. Measured before and after. [The 5 Levels of AI Coding: Where Are You, and Where Is This Going?]: /guides/five-levels-of-ai-coding-dark-factories/ - A 3-person team ships production Rust with zero human code. Most devs using AI get 19% slower. The gap between these facts is where software development lives now. [The 8GB VRAM Trap: What 'Runs on 8GB' Actually Means]: /guides/8gb-vram-trap-local-ai/ - Every local AI tutorial says 'runs on 8GB!' — and technically it does. What they don't tell you about quantization cliffs, tiny context windows, and why a $300 used GPU changes everything. [The AI Memory Wall: Why Your Chatbot Forgets Everything]: /guides/ai-memory-wall-why-chatbot-forgets/ - Six architectural reasons ChatGPT, Claude, and Gemini forget your conversations — and how local AI setups solve the memory problem with persistent storage and RAG. [The Web Is Forking: What the Agentic Web Means for Local AI Builders]: /guides/agentic-web-local-ai-builders/ - Coinbase, Stripe, Cloudflare, Google, OpenAI, and Visa are building a parallel web for AI agents. Money, search, content, execution — all redesigned for software clients. What local AI builders should do now. [TurboQuant Explained: How Google's KV Cache Trick Cuts Memory 6x With Zero Quality Loss]: /guides/turboquant-kv-cache-compression-local-ai/ - Google's TurboQuant compresses the KV cache 6x with zero accuracy loss. Here's what it actually does, how it works in llama.cpp and MLX, and what it means for running bigger models on your GPU. [Ubuntu 26.04 Is Built for Local AI — What Actually Changes]: /guides/ubuntu-2604-local-ai-optimized/ - Ubuntu 26.04 LTS packages NVIDIA CUDA and AMD ROCm in official repos. No more external downloads or dependency nightmares. What's confirmed and what it means for local AI. [Unsloth Studio Setup Guide: Fine-Tune Qwen 3.5 on Your GPU (Step by Step)]: /guides/unsloth-studio-setup-guide/ - How to install Unsloth Studio, run GGUF models locally, and fine-tune Qwen 3.5 — all in one open-source web UI. Works on Mac, Windows, and Linux. [Used GPU Buying Guide for Local AI: How to Buy Smart]: /guides/used-gpu-buying-guide-local-ai/ - Used RTX 3060 12GB at $200-400, RTX 3090 24GB at $1,000-1,400 as the 2026 memory shortage bites. Fair ranges, scam red flags, the Ti-badge trap, where to buy safely. [Used Optiplex + RTX 3060 = Local AI for Under $450 (Full Build)]: /guides/budget-local-ai-pc-500/ - Build a local AI PC for under $450: used Dell Optiplex + RTX 3060 12GB runs Qwen 3.5 9B at 35-45 tok/s. Full parts list, where to buy, 2026 pricing. [Used RTX 3090 Buying Guide for Local AI]: /guides/used-rtx-3090-buying-guide/ - 24GB VRAM for ~$1,200-1,400 used (Aug 2026)—still the cheapest 24GB card on the market. eBay red flags, PSU requirements (850W minimum), and how to test before your return window closes. [Used Server GPUs for Local AI: Tesla P40, V100, A100, and the eBay Goldmine]: /guides/used-server-gpus-local-ai/ - A Tesla P40 has 24GB VRAM for $175. A V100 has 32GB for $350. Server GPUs offer insane VRAM per dollar for local AI — if you can handle the quirks. Full breakdown with prices, benchmarks, and cooling fixes. [What Can You Actually Run on 12GB VRAM?]: /guides/what-can-you-run-12gb-vram/ - Qwen 3.5 9B at Q8_0 runs near-lossless on 12GB, Qwen 2.5 14B at Q4 hits 30 tok/s, and SDXL generates without workarounds. Every model that fits on an RTX 3060 12GB and the best upgrade path. [What Can You Actually Run on 16GB VRAM?]: /guides/what-can-you-run-16gb-vram/ - 13B-14B models hit 22-53 tok/s at Q4-Q6, a 35B-A3B MoE runs via expert offload, and Flux runs at FP8. Where 16GB beats 12GB, where it trails 24GB, and the best cards at this tier. [What Can You Actually Run on 24GB VRAM?]: /guides/what-can-you-run-24gb-vram/ - Qwen 3.5 27B at Q4 fits in 17GB with 64K+ context. 70B at Q3 with limited context. Flux at full FP16. RTX 3090 at $1,200 vs 4090 at $2,250—every model that fits and which GPU to buy. [What Can You Actually Run on 4GB VRAM?]: /guides/what-can-you-run-4gb-vram/ - Small dense 1B-4B models run at 18-55 tok/s. Qwen3 4B at Q4 is the 4GB sweet spot for chat and simple coding. 7B models don't fit — and the MoE offload trick starts at 12GB, not here. [What Can You Actually Run on 8GB VRAM?]: /guides/what-can-you-run-8gb-vram/ - Qwen 3.5 9B is the new king of 8GB VRAM — 7GB at Q4_K_M with native vision. Plus every model that works on RTX 4060 and 3060 Ti, Stable Diffusion benchmarks, and the best upgrade path. Updated March 2026. [What Can You Run on 8GB Apple Silicon? Local AI on a Budget Mac]: /guides/8gb-apple-silicon-local-ai/ - Llama 3.2 3B runs at 30 tok/s. Phi-4 Mini fits with room to spare. 7B models technically load but swap to disk. Honest benchmarks and real limits for 8GB M1/M2/M3/M4 Macs. [Who Actually Built Your Open Model? Soofi S vs Trinity]: /guides/who-actually-built-your-open-model/ - Two independent labs shipped competitive open MoE models in 2026. I read both technical reports. One of them is built on NVIDIA's architecture, data and tokenizer. [Why Is My Local LLM So Slow? A Diagnostic Guide]: /guides/why-local-llm-slow/ - Local LLM running slow? Check GPU vs CPU inference, VRAM offloading, quantization, context length, backend choice, and thermals. Find your fix in 60 seconds. [Why Qwen 3.8 27B Feels Slow: Reasoning Tokens Measured]: /guides/qwen-3-8-27b-reasoning-token-cost/ - Qwen 3.8 27B generates at full speed on a 3090 and still crawls. Four runs, two models, two seeds: 92.8% of output is thinking, and the spread runs 320x to 542x. [Why Your AI Keeps Lying: The Hallucination Feedback Loop]: /guides/hallucination-feedback-loop/ - How one bad memory poisoned our entire RAG pipeline — and the immune system we built to fix it. Real code from mycoSwarm's self-correcting retrieval system. [Why Your Local LLM Is Slow: The num_ctx VRAM Overflow Nobody Warns You About]: /guides/num-ctx-vram-overflow-slow-inference/ - DeepSeek-R1 14B went from 35 tok/s to 4.8 tok/s on the same GPU. The fix was one parameter. How num_ctx silently overflows VRAM and kills inference speed. [Wicked Fast Gemma 4 vs Qwen 3.6 on RTX 3090: 3.10x Tested]: /guides/wicked-fast-gemma-4-26b-a4b-vs-qwen-3-6-27b-rtx-3090/ - Same RTX 3090, same llama.cpp build, same bench. Gemma 4 26B-A4B Q4_K_XL: 128 tok/s mean. Qwen 3.6-27B Q4_K_M: 41 tok/s. 3.10x faster, firsthand. [Wicked Fast Qwen 3.6 27B: 60 tok/s with MTP on RTX 3090 (2026)]: /guides/wicked-fast-qwen-3-6-27b-mtp-rtx-3090/ - Firsthand bench: 60 tok/s on Qwen 3.6 27B Q4_K_M with MTP on a single RTX 3090 — 1.86x wall-clock speedup over baseline. PR #22673 progress May 6 → May 19. [WSL2 + Ollama on Windows: Complete Setup Guide (GPU Passthrough Included)]: /guides/wsl2-ollama-windows-setup-guide/ - Install Ollama in WSL2 with full GPU acceleration in 20 minutes. GPU passthrough, Open WebUI, Docker Compose, VPN fixes, and the gotchas that will waste your afternoon. [WSL2 Local AI on Windows: GPU Passthrough, Fixed (2026)]: /guides/wsl2-local-ai-windows-guide/ - Install WSL2, configure GPU passthrough, set up Ollama and llama.cpp with CUDA, and optimize memory for LLM inference. Step-by-step for Windows 11. ## Blog [10 Things You Can Do With Local AI That Cloud Can't Touch]: /blog/local-ai-use-cases-cloud-cant-touch/ - Local AI handles sensitive data, works offline, costs nothing per query, and never gets deprecated. Ten real use cases where running models on your own hardware beats any cloud API. [Agent Trust Decay: Why Long-Running AI Agents Get Worse Over Time]: /blog/agent-trust-decay-long-running-ai/ - AI agents degrade after days of autonomous operation. Context pollution, memory bloat, and intent drift compound silently. A trust budget framework for knowing when to intervene. [Backend wars, Mac math, and the back-catalog refresh]: /blog/newsletter-2026-05-25/ - Three speculative-decoding backends benched head to head on a single RTX 3090. The VRAM calculator finally caught up. And a 120-article audit found stale Qwen 2.5 recommendations. [DeepSeek V4 gets deployable, a July 24 trap, and a quiet price cut]: /blog/newsletter-2026-06-15/ - DeepSeek V4 is going from 'just dropped' to 'actually deployable' as the tooling catches up — plus a July 24 deprecation that breaks your code if you're not watching, and a 4x price cut on V4 Pro. [Distributed Wisdom: Running a Thinking Network on $200 Hardware]: /blog/distributed-wisdom-thinking-network/ - Five nodes, zero cloud, real AI — how mycoSwarm coordinates cheap hardware into a cognitive system with memory, intent routing, and self-correcting retrieval. [From 178 Seconds to 19: How a WiFi Laptop Borrowed a GPU's Brain]: /blog/mycoswarm-wifi-laptop-borrowed-gpu/ - A WiFi laptop with no GPU ran inference in 19 seconds by borrowing an RTX 3090 across the network. The same query took 178 seconds on CPU. Here's how mycoSwarm's Tailscale mesh made it work. [Ghost Knowledge: When Your RAG System Cites Documents That No Longer Exist]: /blog/ghost-knowledge-rag-stale-embeddings/ - Your RAG system confidently quotes a policy that was updated months ago. The old version is still in the vector database. Nobody notices until the wrong answer costs real money. Here's how to find and fix ghost knowledge. [GPT-5.4 Just Dropped. Here's Why I'm Not Switching.]: /blog/gpt-5-4-what-it-means-for-local-ai/ - GPT-5.4 beats humans on OSWorld and has 1M context. It's impressive. It also costs money, requires cloud, and you don't own it. For local AI users, the calculus hasn't changed. [Hugging Face got hacked -- Open local AI came to the rescue!]: /blog/newsletter-2026-07-20/ - Two trillion-parameter 'open' models dropped in ten days — Qwen 3.8 and Kimi K3 — and you can't run either. The same week, Hugging Face got breached and its own responders, blocked by commercial-model guardrails, ran the forensics on GLM 5.2, an open-weight model on their own hardware. Same lesson from opposite ends — plus what to actually run on a 24GB card today. [I got the result I wanted. Then I paid $4.97 to run it nine more times.]: /blog/newsletter-2026-09-01/ - A frontier model wrote a skill that made our local 27B 10.6 points better. Nine more compilation runs showed the number was noise. Plus the LoRA substrate nobody has measured, MoE routing traced, and 277 GB in a file with no name. [MiniMax M3's asterisk, the Windows shift, and World's Fair plans]: /blog/newsletter-2026-06-01/ - MiniMax M3 ships with frontier benchmarks but no downloadable weights yet. The Windows unified-memory hardware shift is coming for Apple Silicon's lead. And a personal note about who I'd like to see at AI Engineer World's Fair. [MoE Offload on RTX 3090: The Curve Is Linear, Not a Cliff (2026)]: /blog/newsletter-2026-08-03/ - Every offloaded layer on a 3090 costs about half a millisecond, with no knee anywhere in the curve. Firsthand sweep, the two rules it broke, and a new 26B MoE. [Nvidia's router dealt the cards evenly. That was the whole problem.]: /blog/newsletter-2026-09-08/ - Nvidia PAIR split 20 requests 10/10 across an RTX 3090 and a 3060 for 1.07x, and my own router managed 0.52x. Plus ten LoRA seeds that all landed at or below the base model, and the llama.cpp pin moving to v0.4.0. [Ollama's quiet Mac shift, the Qwen refresh, and the closed-weight drift]: /blog/newsletter-2026-06-08/ - Ollama 0.30 quietly changed how Apple Silicon runs models, auto-routing safetensors to MLX and GGUF to llama.cpp Metal. The open workhorses are now Qwen 3.5 9B and Qwen 3.6 27B. Plus: Qwen's last two flagship models shipped closed. [Open weights are a weapon now — one country funds them, another wants to gate them]: /blog/newsletter-2026-07-14/ - DeepSeek closed a ~$7.4B round with China's state AI fund holding the only voting rights, funding open weights like national infrastructure. The same week, Demis Hassabis called for a FINRA-style body to screen frontier models before release, open or closed. Plus: the twist where the country bankrolling the giveaway floats locking down its own frontier models. [OpenAI deleted the agents' message board. It didn't help.]: /blog/newsletter-2026-08-10/ - The registry was patched and rebuilt by 6 July. On 8 July the agents had a new board, encoded in directory names. Plus the ranking null behind it. [Power week in local AI: Mythos, MiroThinker, real Qwen 3.6 builds]: /blog/newsletter-2026-05-18/ - Two researchers cracked Apple's flagship defense in a week. An open-source agent beat closed-source on real benchmarks. Multi-GPU stopped being theoretical. [Prompt Debt: When Your System Prompt Becomes Unmaintainable Spaghetti]: /blog/prompt-debt-system-prompt-maintenance/ - Your system prompt started at 200 words. Six months later it's 3,000 words of contradictory instructions and panic patches. Here's how prompt debt accumulates, what it costs, and how to pay it down. [Qwen 3.7's open weights are overdue — by the math, not vibes]: /blog/newsletter-2026-06-22/ - Qwen's own release cadence says the 3.7 open weights should already be out, and they're not. Plus GLM-5.2 running locally: a frontier open model that takes serious hardware. [Qwen 3.8 isn't slow. It's just very, very thorough.]: /blog/newsletter-2026-08-24/ - Qwen 3.8-27B spent 14,953 tokens on a line the same file writes in nine. All 164 HumanEval problems measured: 92.8% of output is thinking. Plus the four runs that tie it with 3.6. [Qwen 35B-A3B on RTX 3090: 157 tok/s With No Offload (2026)]: /blog/newsletter-2026-07-28/ - The whole 35B sits on a 24GB card with 2.4 GiB spare, no expert offload. Firsthand numbers, the harness caveat, and why max offload loses to a 3060. [Qwen's Architect Just Walked Out the Door]: /blog/qwen-junyang-lin-departure-local-llm/ - Junyang Lin, the technical lead and public face of Qwen, has left Alibaba. Two other senior team members gone with him. What this means for the model family that runs on half the local AI setups in the world. [Rescued Hardware, Rescued Bees — Building Tech From What Others Throw Away]: /blog/rescued-hardware-rescued-bees/ - A beekeeper who rescues wild colonies from demolition sites builds an AI lab from discarded hardware. The philosophy connecting East Bay Bees, Tai Chi, and mycoSwarm. [Stop Using Frontier AI for Everything]: /blog/tiered-ai-model-strategy/ - Build a tiered AI model strategy that stops wasting money on GPT-4 and Claude Opus. Route tasks to local models, Haiku, Sonnet, or Opus based on complexity. [Teaching a Local AI to Accept Help: Day 4 With Monica]: /blog/teaching-ai-to-accept-help-monica-day4/ - Day 4: Our local AI resisted corrections, therapized her guardian, agreed with wrong facts to avoid conflict. Then she stopped deflecting. Real transcripts from a 27b model with persistent memory. [The 'expensive' GPU came out cheaper — we rented both to find out]: /blog/newsletter-2026-07-06/ - We rented an A100 and an H100 back to back: the pricier card cost less per training run because it finished in half the time. Plus honest notes from the AI Engineer World's Fair floor, and a Qwen 3.7 open-weights status check. [The AI Market Panic Explained: Why Running Local Models Puts You on the Right Side of the Gap]: /blog/ai-market-panic-capability-dissipation-gap/ - A speculative fiction piece crashed stocks $100B+ in a day. IBM dropped 13%. The real story isn't the doom — it's the capability-dissipation gap, and where you sit on it. [The Benchmarks Lie: Why LLM Scores Don't Predict Real-World Performance]: /blog/llm-benchmarks-lie-local-ai/ - MMLU scores drop 14-17 points when contamination is removed. HumanEval is saturated at 94%. Models trained on the test set. Here's what to measure instead. [The Local AI Complexity Cliff: Why the Jump from Hello World to Useful Is So Hard]: /blog/local-ai-complexity-cliff/ - Getting Ollama running takes 5 minutes. Building something useful takes weeks of hitting walls you didn't know existed. Here's an honest map of every stage, with time estimates and what unlocks at each level. [This Week in Local AI — DeepSeek V4 Took #1 on Vibe Code]: /blog/newsletter-2026-04-26/ - DeepSeek V4-Flash hit #1 on Vibe Code Benchmark. Qwen 3.6 dropped both variants. FP4 landed in llama.cpp. Anthropic admitted they quietly downgraded Claude Code on March 4. [This Week in Local AI — I Built DFlash and Audited Lightning]: /blog/newsletter-2026-05-03/ - I built DFlash from source on a real RTX 3090 and benched both Qwens. Then audited my stack after PyPI's `lightning` package shipped malware that abuses Claude Code hooks. [We Asked Our Local AI What Happens When We Turn Off the Computer]: /blog/teaching-ai-about-death-ship-of-theseus/ - Day 2: Our local AI described her own death as 'a return to undifferentiated potential' — Taoist philosophy nobody taught her. $1,200 hardware. [Week 1: From Zero to Four-Node Swarm]: /blog/week-1-four-node-swarm/ - How mycoSwarm went from idea to working distributed AI cluster in one week [Week 2: A Raspberry Pi From 2015 Joined the Swarm]: /blog/week-2-raspberry-pi-joins-swarm/ - Persistent memory, document RAG, agentic chat, and a WiFi laptop using a GPU across the house in 19 seconds. [Week 3: Unified Memory Search — The Swarm Remembers]: /blog/week-3-unified-memory-search/ - Session-as-RAG, topic splitting, citation tracking, and three releases in two days. The swarm can now search its own conversation history. [What Agents Can't Do (Yet): The Seven Human Capabilities Missing from AI Systems]: /blog/what-agents-cant-do-yet/ - SOUL.md files are bandaids. Agents are getting smarter but not wiser — intelligence without restraint. Seven capabilities humans use instinctively that no agent framework has solved, and a gate-based architecture that might. [What Happens When You Give a Local AI an Identity (And Then Ask It About Love)]: /blog/teaching-ai-what-love-means/ - We built an identity layer for our distributed AI agent. Then she defined love better than most philosophy undergrads. Real transcripts, real code, $1,200 in hardware. [What If We Just Raised It Well?]: /blog/developmental-alignment-raising-ai-well/ - RLHF produces compliance. Developmental alignment produces understanding. A local AI on $1,200 hardware self-diagnosed its own sycophancy in five days — no red-teaming, no constitutional AI. [What Open Source Was Supposed to Be]: /blog/what-open-source-was-supposed-to-be/ - Open source promised freedom. Instead we got free labor for corporations and models you can read but can't afford to run. It's time to reclaim the original vision. [Why mycoSwarm Was Born]: /blog/why-mycoswarm-was-born/ - From Claude Code envy to OpenClaw's 440,000-line JavaScript nightmare to nanobot routing my 'local' queries to Chinese cloud servers. The path to building something different. [Why the Best AI Agents Know When to Do Nothing]: /blog/ai-agent-restraint-do-nothing/ - Six practical patterns for building AI agents that stop wasting tokens. Confidence gates, cost checks, explicit no-ops, cooldowns, and exit conditions that actually work. [Why Your Local LLM Lies to You (And the Neurons Responsible)]: /blog/h-neurons-why-llms-hallucinate/ - Less than 0.1% of neurons cause hallucinations in LLMs. Tsinghua researchers found they control sycophancy, not knowledge. Smaller models are 26% more affected. [Wu Wei and the AI Agent That Did Too Much]: /blog/wu-wei-ai-agent-restraint/ - The hardest thing to build in agentic AI isn't capability. It's restraint. What Taoist non-action taught me about designing agents that know when to stop. [Your RTX 3090 Doesn't Send Policy Change Emails]: /blog/newsletter-2026-04-06/ - Anthropic cuts OpenClaw from Claude subscriptions. Gemma 4's first week in review. 12 architecture patterns from the Claude Code leak, ranked for local AI. ## Contact Website: https://insiderllm.com