Structured Output
What Is Jev, and Can You Run One on Your Own GPU?
Jev picks from options you supply instead of writing, in one pass, with a probability per answer. What it is, what it costs, and the open version for a 3090.
Stop writing the answer. Pick it. 47 items waiting.
Codacus's parallel-decision llama.cpp branch answers a schema in one forward pass, and the ten-seeds 47-item split is the ready-made test. Plus ChatGPT-User fetches sliding 5 percent a week while OAI-SearchBot tripled, the v0.4.0 pin on the 3060, and a third bench rig.
Jev Mode on a 3090: 26 ms per Token You Don't Write
Jev-style parallel decisions in llama.cpp on Qwen3.6-27B, RTX 3090: 1.3x faster at four tokens, 5.5x at fifty, 21 vs 23 of 47. The gain is the tokens you skip.
Best Local LLMs for Structured Output: Qwen 3.6, Gemma 4
JSON schema, grammar constraints, and Outlines compared. Current model picks: Qwen 3.6, Gemma 4, DeepSeek V4. Common failures + working code. May 2026.
A weekly email with every new guide and measured benchmark.
Subscribe — free, no spam