NVFP4
FP4 Just Landed in llama.cpp: NVFP4 vs MXFP4 Explained (2026)
NVFP4 in llama.cpp, MXFP4 in ik_llama.cpp. The first practical FP4 quantization for the GGUF ecosystem — what works, what doesn't, and what to test.
llama.cpp Build Errors: Common Fixes for Every Platform
llama.cpp won't build or runs wrong? CMake, CUDA, Gemma 4 thinking-mode, Qwen 3.6 kwargs, num_ctx VRAM overflow. Exact fixes for every platform.
Best Local FLUX Setup: FLUX.2, FLUX.1, RTX 3090 (2026)
FLUX.1 vs FLUX.2: 12GB VRAM minimum with GGUF, 24GB at full quality. ComfyUI / Forge / Python setup, May 2026 hardware, NVFP4 on Blackwell.