Models Guides
48 InsiderLLM guides in Models — practical, tested walkthroughs for running AI locally, sorted by most recently updated.
- MoE Models Explained: Why Mixtral Uses 46B Parameters But Runs Like 13B Mixture of Experts explained for local AI — why MoE models run fast but still need full VRAM. Mixtral, DeepSeek V3, DBRX compared with dense model alternatives.
- Best Way to Run Qwen 3.6 35B MoE Locally: VRAM, Speed, Setup Qwen 3.6-35B-A3B has 35B total params but only 3B active per token. Real tok/s on RTX 3090, 4090, 5070 Ti, dual 5060 Ti, and M3 Ultra. Quants and setup.
- Kimi K3 & Qwen 3.8: Open Weights You Can't Run (2026) Kimi K3's 2.8T weights need 64 accelerators to load. Qwen 3.8's 27B did ship, Apache 2.0, and fits 24GB. Openness and runnability are separate axes.
- Who Actually Built Your Open Model? Soofi S vs Trinity Two independent labs shipped competitive open MoE models in 2026. I read both technical reports. One of them is built on NVIDIA's architecture, data and tokenizer.
- Qwen 3.6 Complete Guide: 27B Dense, 35B-A3B MoE, and Which to Use Qwen 3.6 landed in two open-weight flavors: 27B dense and 35B-A3B MoE. Benchmarks, hardware fit, and which variant to run on your GPU.
- Qwen 3.6 MoE Routing, Measured: Flat Is the Wrong Number I traced every expert routing decision Qwen 3.6-35B-A3B makes across six workloads on an RTX 3060. Routing isn't flat, and 112 slots is the whole answer.
- Best Local Coding Models Ranked: Every VRAM Tier, Every Benchmark (2026) The best local LLMs for coding in 2026, ranked by VRAM tier. Qwen 3.6-27B, 3.6-35B-A3B, DeepSeek V4-Flash, benchmarks, editor setup, and Claude Code alternatives.
- How to Run GLM 5.2 Locally: GPU, VRAM & Quant Guide GLM 5.2 is 753B params and 1.51TB at full precision. Run it locally: the live Unsloth quant ladder, every GPU and RAM path, and the quant to actually target.
- Best Qwen Models Ranked: Which to Run Locally (Mid-2026) Every Qwen you can run locally: 3.8-27B (new, Apache 2.0), 3.6, 3.5, Qwen3-Coder-Next, Qwen-VL. VRAM per tier, Ollama setup, and how they bench.
- The Benchmarks Lie: Why LLM Scores Don't Predict Real-World Performance MMLU scores drop 14-17 points when contamination is removed. HumanEval is saturated at 94%. Models trained on the test set. Here's what to measure instead.
- Inkling 975B vs Your 3090: The Real Memory Math (2026) Inkling's smallest 1-bit quant is 270GB. A maxed consumer desktop holds 256GB. The open frontier left consumer hardware behind. Here's the honest math.
- Best Local LLMs for Writing & Creative Work Llama 3.3 70B is the best local prose model in 2026; Qwen3 32B is the 24GB sweet spot for fiction and long-form. Model picks for every VRAM tier and writing task.
- CodeLlama vs DeepSeek Coder vs Qwen Coder: Best Local Coding Models Compared CodeLlama vs DeepSeek Coder vs Qwen Coder vs Codestral benchmarked: HumanEval scores, VRAM per quant, and speed tests. Qwen 7B beats CodeLlama 70B.
- Model Formats Explained: GGUF vs GPTQ vs AWQ vs EXL2 GGUF vs GPTQ vs AWQ vs EXL2/EXL3 model formats explained — plus the new FP4 (MXFP4/NVFP4) and Apple's MLX. What each does, which tools run it, and how to choose for your GPU.
- Best Local LLMs for Math & Reasoning: What Actually Works The best local LLMs for math and reasoning in 2026, ranked by VRAM tier. AIME 2026 and GPQA benchmarks for Qwen 3.6, Qwen 3.5 thinking, and where the old R1-distills now stand.
- Mixtral 8x7B & 8x22B VRAM Requirements Mixtral 8x7B and 8x22B VRAM requirements at every quantization level — plus why the 'every expert must sit in VRAM' rule Mixtral taught no longer holds in the A3B era.
- Hugging Face Hacked by AI Agent — Saved by a Local Model (2026) Hugging Face disclosed July 16 that an autonomous AI agent breached its internal infra — now confirmed as OpenAI's own models. The models you download are safe — here's what was and wasn't hit.
- Llama 3 Guide: Every Size from 1B to 405B Complete Llama 3 guide covering every model from 1B to 405B. VRAM requirements, Ollama setup, benchmarks vs Qwen 3, and which size fits your hardware.
- DeepSeek Models Guide: R1, V3, and Coder Complete DeepSeek models guide covering R1, V3, and Coder locally. Which distilled R1 to pick for your GPU, VRAM requirements, and benchmarks vs Qwen 3.
- Qwen3 Complete Guide: Every Model from 0.6B to 235B Qwen3 is the best open model family for budget local AI. Dense models from 0.6B to 32B, MoE models that punch above their weight, and a /think toggle no one else has.
- DeepSeek V4 Flash vs Pro: Verdict, Cost, Setup (2026) Flash or Pro? Which DeepSeek V4 to run, what each costs (Haiku-tier pricing), and how to deploy locally. The verdict, the tradeoffs, no launch recap.
- Phi Models Guide: Microsoft's Small but Mighty LLMs Phi-4 14B scores 84.8% on MMLU — matching models 5x its size — and fits on a 12GB GPU at Q4. The full Phi-4 lineup, including the reasoning and vision-reasoning variants, with VRAM needs, benchmarks, and honest weaknesses.
- Best Qwen 3.5 Setup: When to Stay vs Move to 3.6 (2026) 3.5 is Qwen's stable open workhorse — 3.6 replaced only two tiers, 3.7 went closed. Which 3.5 model on which GPU, when to stay vs move to 3.6.
- Is Qwen Going Closed? Open Weights vs Frontier (2026) Qwen split into a closed frontier (Max, Plus, VLA) and an open mid-tier (3.6-27B and 35B-A3B). The 3.7 open weights aren't here yet. The honest read.
- DeepSeek V3.2 Guide: What Changed and How to Run It Locally DeepSeek V3.2 was the Feb 2026 flagship — V4 now leads. But the R1-Distill models run on a $200 used GPU and remain the local reasoning pick.
- Qwen 3.5 Locally — 27B vs 35B-A3B vs 122B, Which Model Fits Your GPU Qwen 3.5 and 3.6 on local hardware. 27B dense vs 35B-A3B MoE vs 122B compared. VRAM tables, community tok/s on RTX 3090, and which to pick for your card.
- Llama 4 Guide: Running Scout and Maverick Locally (2026) Complete Llama 4 Scout (109B MoE) and Maverick guide for local AI. VRAM, Ollama and vLLM setup, hardware reality, and how it stacks against Qwen 3.6.
- Best 8GB GPU Model: How to Set Up Qwen 3.5 9B (Step by Step) Qwen 3.5 9B fits in 6.6GB and beats Qwen 3-class models 3x its size. Setup on Ollama/llama.cpp, quant table, where 9B still fits in the May 2026 lineup.
- Gemma Models Guide: Google's Lightweight Local LLMs Gemma 3 27B beats Gemini 1.5 Pro on benchmarks and runs on a single GPU. The 4B outperforms Gemma 2 27B. Full lineup from 1B to 27B with VRAM needs, speeds, and honest comparisons.
- Llama 4 vs Qwen3 vs DeepSeek V3.2: Which to Run Locally in 2026 Llama 4 needs 55GB. DeepSeek V3.2 needs 350GB. Qwen3 runs on 8GB. Here's who wins at each VRAM tier and use case for local AI in 2026.
- Are Mistral Models Still Worth Running? Only Nemo 12B (Here's Why) Mistral Medium 3.5-128B dropped April 29, 2026: dense 128B, 256k context, Modified MIT. Hardware reality, license caveats, which Mistral to actually run.
- Best Way to Run Qwen 3.5 on Mac: MLX vs Ollama Speed Test MLX runs Qwen 3.5 up to 2x faster than Ollama on Apple Silicon. Head-to-head benchmarks on M1 through M4, with setup instructions for both.
- Best Qwen 3.5 Models Ranked: Every Size, Every GPU, Every Quant Complete ranking of all Qwen 3.5 models from 0.8B to 397B. VRAM requirements, speed benchmarks, and which model to pick for your hardware.
- Gemma 4 Just Dropped: What Local AI Builders Need to Know Google's Gemma 4 is here -- dense and MoE variants, Apache 2.0, multimodal with vision and audio. VRAM requirements, benchmarks, and how it compares to Qwen 3.5.
- Best Local LLMs for Chat & Conversation The best local LLMs for chat and conversation in 2026. Picks for every VRAM tier from 8GB to 24GB, with Ollama commands to start chatting immediately.
- Quantization Explained: What It Means for Local AI Q4_K_M shrinks a 7B model from 14GB to ~4GB while keeping 90-95% quality. What every quantization format means, how much VRAM each saves, and which to pick for your GPU.
- Mistral Voxtral TTS: Open-Weight Voice AI You Can Run Locally Voxtral TTS is a 4B open-weight text-to-speech model that beats ElevenLabs Flash v2.5 in blind tests. 70ms latency, 9 languages, voice cloning from 3 seconds. Here's how to run it.
- Best Models Under 3B: Small LLMs That Work The best models under 3B parameters for laptops, old GPUs, Raspberry Pi, and phones. What works, what doesn't, and which tiny LLM to pick for your use case.
- RWKV-7: Infinite Context, Zero KV Cache — The Local-First Architecture RWKV-7 uses O(1) memory per token. Context length doesn't increase VRAM. At all. 16 tok/s on a Raspberry Pi. Here's why it matters for local AI and how to run it.
- Qwen 3.5 Small Models: The 9B Beats Last-Gen 30B — Here's What Matters for Local AI Alibaba's Qwen 3.5 drops 4 small models (0.8B to 9B) — all natively multimodal, 262K context, Apache 2.0. The 9B beats Qwen3-30B on reasoning and destroys GPT-5-Nano on vision. VRAM tables and what to run.
- DeepSeek V4: Everything We Know Before It Drops DeepSeek V4 launches next week with native image and video generation, 1M context, and rumored 1T MoE params with only 32B active. Here's what local AI builders need to know and how to prepare.
- LiquidAI LFM2: The First Hybrid Model Built for Your Hardware LFM2-24B-A2B runs at 112 tok/s on CPU with only 2.3B active params. Not a transformer. GGUF files from 13.5GB, Ollama and llama.cpp setup, and where it beats Qwen.
- Distilled vs Frontier Models for Local AI — What You're Actually Getting That local model you love was probably trained on stolen outputs from Claude or GPT. Here's what distillation actually does to a model's reasoning, where it breaks, and why it matters most for agentic work.
- nanollama: Train Your Own Llama 3 From Scratch on Custom Data Pretrain Llama 3 architecture models from raw text, export to GGUF, and run with llama.cpp. Forked from Karpathy's nanochat. 46M to 7B parameters.
- Qwen vs Llama vs Mistral: Which Model Family Should You Build On? Qwen has 201 languages and a model for every task. Llama has the biggest community. Mistral pioneered efficient MoE. Decision framework for choosing your model family in 2026.
- Ouro-2.6B-Thinking: ByteDance's Looped Model That Punches Like an 8B Ouro-2.6B loops through the same transformer blocks 4 times to match 8B models at 2.6B parameters. Under 2GB at Q4. How the architecture works and why it matters.
- Mixtral VRAM Requirements: 8x7B and 8x22B at Every Quantization Level Mixtral 8x7B has 46.7B params but only 12.9B activate per token. You still need VRAM for all 46.7B. Exact VRAM for every quant from Q2 to FP16.
- GPT-OSS Guide: OpenAI's First Open Model for Local AI GPT-OSS 20B is OpenAI's first open-weight model. MoE with 3.6B active params, MXFP4 at 13GB, 128K context, Apache 2.0. Here's how to run it.