Inference Guides
1 InsiderLLM guides in Inference — practical, tested walkthroughs for running AI locally, sorted by most recently updated.
- Flash-MoE: What a 397B Model on a Laptop Actually Cost Flash-MoE ran Qwen3.5-397B on a 48GB MacBook at 4.4 tok/s — then stopped dead in March 2026, unlicensed. The 2-bit trap it documented is the part worth keeping.