Llm-Inference
Best Way to Run 31B Models on a Laptop? Treat Them Like Databases
LARQL decompiles transformer weights into a queryable graph called a vindex. The project pitches a new shape for local inference: walk a subgraph, patch facts, stream from disk. Here's what's real, what's claimed, and what's still research.
Flash-MoE: What a 397B Model on a Laptop Actually Cost
Flash-MoE ran Qwen3.5-397B on a 48GB MacBook at 4.4 tok/s — then stopped dead in March 2026, unlicensed. The 2-bit trap it documented is the part worth keeping.
LocalAgent: A Local-First Agent Runtime That Actually Cares About Safety
Rust CLI for AI agents with deny-by-default permissions, approval workflows, and deterministic replay. Works with LM Studio, Ollama, and llama.cpp.
Best Dual-GPU Local AI Setup: RTX 3090, 5060 Ti (2026)
Dual RTX 3090, 2x RTX 5060 Ti, 2x 2080 Ti modded, mixed setups: real configs for Qwen 3.6, MoE, 70B. Tensor vs pipeline parallelism, llama.cpp/vLLM.
A weekly email with every new guide and measured benchmark.
Subscribe — free, no spam