Apple Silicon
Gemma 4 26B in 2GB RAM: The MoE Memory Ladder Explained
One model, three places its experts can live: VRAM, RAM, SSD. We measured the first two on Gemma 4. TurboFieldfare just added the third.
Ollama's quiet Mac shift, the Qwen refresh, and the closed-weight drift
Ollama 0.30 quietly changed how Apple Silicon runs models, auto-routing safetensors to MLX and GGUF to llama.cpp Metal. The open workhorses are now Qwen 3.5 9B and Qwen 3.6 27B. Plus: Qwen's last two flagship models shipped closed.
Mistral Voxtral TTS: Open-Weight Voice AI You Can Run Locally
Voxtral TTS is a 4B open-weight text-to-speech model that beats ElevenLabs Flash v2.5 in blind tests. 70ms latency, 9 languages, voice cloning from 3 seconds. Here's how to run it.
Flash-MoE: Run a 397B Model on a 48GB Laptop (Here's How)
Flash-MoE streams Qwen3.5-397B from your SSD at 4.4 tok/s using 5.5GB of RAM. Pure C and Metal, no Python. Here's what's real, what's hype, and how to try it.
Apple Neural Engine for LLM Inference: What Actually Works
Apple Silicon has a dedicated Neural Engine that most LLM tools ignore. Here's what it can do for inference, what it can't, and whether ANE-based tools like ANEMLL are worth trying today.
Best Apple M5 Pro and Max for Local AI (2026)
M5 Pro at 307GB/s, M5 Max at 614GB/s (or 460 on the 32-core bin), up to 128GB — now the highest-memory Mac you can buy. Picks for Qwen 3.6 and Llama 3.3 70B.
OpenClaw on Mac: Setup, Optimization, and What Actually Works
brew install openclaw-cli, connect Ollama, configure the gateway, and stop fighting macOS. Apple Silicon setup, memory math, launchd config, and the gotchas nobody warns you about.
What Can You Run on 8GB Apple Silicon? Local AI on a Budget Mac
Llama 3.2 3B runs at 30 tok/s. Phi-4 Mini fits with room to spare. 7B models technically load but swap to disk. Honest benchmarks and real limits for 8GB M1/M2/M3/M4 Macs.
Stable Diffusion on Mac: Image Generation with MLX and Draw Things
Draw Things generates SD 1.5 images in 8-15 seconds on an M2 Pro. ComfyUI takes 3x longer. MLX is fastest but code-only. Complete Mac image gen guide with speed tests.
Ollama on Mac: Setup and Optimization Guide (2026)
Install Ollama on Apple Silicon, verify Metal GPU is active, and tune it for your Mac's RAM. Config for M1 through M4 Ultra with model picks per memory tier.
Ollama on Mac Not Working? Fix Metal, Memory Pressure, and Slow Performance
ollama ps says CPU? Generation crawling at 2 tok/s? macOS killed your model mid-sentence? Every Mac-specific Ollama problem diagnosed and fixed with exact commands.
Mac Studio for Local AI: Is It Worth the Price?
Mac Studio M4 Max (64GB) and M3 Ultra (96GB) for local LLMs after Apple's 2026 memory cuts. Real tok/s, cost vs dual RTX 3090, and who should buy one.
LM Studio vs Ollama on Mac: Which Should You Use?
LM Studio's MLX backend is 20-30% faster and uses half the memory. Ollama is lighter, always-on, and better for APIs. Mac-specific benchmarks and when to use each.
Fine-Tuning on Mac: LoRA & QLoRA with MLX
Fine-tune Llama, Qwen, and Mistral on Apple Silicon using mlx-lm. Real memory numbers, step-by-step commands, and how to deploy your model with Ollama.
Best Way to Run Qwen 3.5 on Mac: MLX vs Ollama Speed Test
MLX runs Qwen 3.5 up to 2x faster than Ollama on Apple Silicon. Head-to-head benchmarks on M1 through M4, with setup instructions for both.
M4 Max and M3 Ultra for Local LLMs: Apple Silicon in 2026
No M4 Ultra exists. After Apple's 2026 memory cuts, the Mac Studio pairs the M4 Max (64GB) with the M3 Ultra (96GB, 819 GB/s). Which to buy for local AI.
Mac Mini M4 for Local AI: Which Config to Buy and What It Actually Runs
Mac Mini M4 Pro 48GB runs Qwen 3.6-35B-A3B silently at 40W. Which config to buy after Apple's 2026 price hikes, and what each tier actually runs for local AI.
Best Local LLMs for Mac in 2026 — M1 through M5 Tested
Best model for every Mac tier, 8GB to the 96GB Studio ceiling. Qwen 3.6, Llama 4 Scout, DeepSeek V4, MLX vs Ollama. Why bandwidth, not RAM, sets your speed.
Run LLMs on Mac M-Series: Faster, Without the Gotchas (2026)
Foundational how-to for Apple Silicon local AI: unified memory, MLX vs Ollama vs llama.cpp Metal, verification, and the headless Mac Mini AI server.
Laptop vs Desktop for Local AI: Which Should You Buy?
A $1,200 desktop RTX 3090 gives you 24GB VRAM. The same money in a gaming laptop gets 8GB. MacBooks break the rules with 48GB+ unified memory for 70B models.
Mac vs PC for Local AI: Which Should You Choose?
An RTX 3090 runs 7B-32B models 2-3x faster than a Mac. A 96GB Mac Studio or a Strix Halo mini-PC (from ~$1,499) loads 70B. Benchmarks, current 2026 prices, and which platform fits.
Best VRAM Cheat Sheet for Local LLMs: Every Model, Every Quant
Exact VRAM for Qwen 3.6, Qwen 3.5, Llama, Mistral, and DeepSeek at Q3 through FP16. Lookup tables for 7B, 9B, 13B, 27B, 32B, 70B, and 120B models with real measurements and GPU recommendations. Updated July 2026.