Guides

249 practical guides for running AI locally — from first install to advanced optimization.

Recently Updated

Getting Started (8)

All 8 Getting Started guides →

Hardware & GPUs (51)

All 51 Hardware & GPUs guides →

Mac & Apple Silicon (14)

All 14 Mac & Apple Silicon guides →

Image Generation (11)

All 11 Image Generation guides →

Models (45)

All 45 Models guides →

Software & Tools (28)

  • Local LLMs vs ChatGPT: An Honest Comparison ChatGPT has web search, voice mode, and GPT-5.2. Local LLMs have privacy, no subscriptions, and no rate limits. Here's when each one wins, what the cost math actually looks like, and why most power users run both.
  • llama.cpp vs Ollama vs vLLM: One User vs Many (2026) Single-user, the three are closer than benchmark posts admit. Concurrent, vLLM pulls 10-20x ahead. Decision tree, the vLLM VRAM gotcha, mid-2026 versions.
  • Run Your First Local LLM in 15 Minutes Install Ollama, pull a model, and chat with AI offline—all in 15 minutes. Works on any Mac, Windows, or Linux machine with 8GB RAM. No accounts, no API keys, no fees.
  • Text Generation WebUI Setup Guide (2026) Install oobabooga's TextGen (formerly text-generation-webui), load GGUF/EXL2/EXL3 models, and configure GPU offloading. Now with vision, tool-calling, and an Anthropic-compatible API. Covers the settings most guides skip.
  • How to Fix Slow Qwen 3.6 27B on RTX 3090 (10-80 tok/s) Qwen 3.6-27B at 12 tok/s on a 3090 when others report 35? The 8-step diagnostic checklist for offload, quants, templates, power limits, and backend choice.
All 28 Software & Tools guides →

AI Agents & OpenClaw (46)

All 46 AI Agents & OpenClaw guides →

Use Cases (41)

  • Local LLMs vs ChatGPT: An Honest Comparison ChatGPT has web search, voice mode, and GPT-5.2. Local LLMs have privacy, no subscriptions, and no rate limits. Here's when each one wins, what the cost math actually looks like, and why most power users run both.
  • Stable Diffusion Locally: Getting Started SD 1.5 runs on 4GB VRAM, SDXL needs 8GB, Flux needs 12GB+. Generate unlimited images for free in under 5 minutes with Fooocus or ComfyUI. Setup, models, and first image tips.
  • Local AI for Privacy: What's Actually Private Running AI locally keeps prompts off corporate servers — but model downloads, telemetry, and VS Code extensions can still leak data. Here's what's genuinely private, what isn't, and how to close every gap.
  • Fine-Tuning LLMs on Consumer Hardware: LoRA and QLoRA Guide Fine-tune a 7B model on 6-10GB VRAM with QLoRA and Unsloth (2-5x faster, 70% less memory). Only 200-500 examples needed — and in 2026 you can fine-tune a 30B MoE on a single 24GB card. Dataset prep through training on RTX 3060-5090.
  • Best Local LLMs for Translation: What Actually Works NLLB handles 200 languages on 3GB VRAM. Qwen 3 and Gemma 3 rival DeepL for European pairs. Opus-MT runs at 300MB per direction. Which local translation model fits your hardware and language needs.
All 41 Use Cases guides →

Architecture & Theory (16)

All 16 Architecture & Theory guides →

Troubleshooting (20)

All 20 Troubleshooting guides →