๐Ÿ“š More on this topic: Best Uncensored Local LLMs ยท Best Models for Coding ยท Best Models Under 3B ยท VRAM Requirements

Cloud AI writes well, but it reads everything you write. Your novel drafts, journal entries, client work, half-formed ideas โ€” all stored on someone else’s servers. Local models let you write, brainstorm, edit, and experiment without sending a single word to the cloud.

The catch: not every local model writes well. Some produce generic, stilted prose. Others refuse to write conflict, romance, or anything remotely dark. And the difference between a 7B and a 32B model for writing quality is enormous โ€” far bigger than for coding or Q&A tasks.

This guide covers which models actually produce good writing, organized by what you want to write and what hardware you have.


What Makes a Good Writing Model?

Writing is one of the hardest tasks for local LLMs. Unlike coding (where output is correct or isn’t) or Q&A (where facts are verifiable), good writing requires:

  • Coherence over long passages โ€” maintaining tone, character, and narrative threads
  • Stylistic range โ€” matching different voices, genres, and registers
  • Instruction following โ€” doing what you ask without drifting
  • Not being boring โ€” avoiding the same safe, generic, corporate-sounding prose

Bigger models are significantly better at all of these. The quality jump from 14B to 32B for writing is more dramatic than for almost any other task. If you can run 32B, do it.


Best Models by VRAM Tier

VRAMModelQuantBest ForQuality
8 GBQwen3-8BQ4_K_MFiction, structured draftsBest all-round at this size
8 GBLlama 3.1 8B InstructQ4_K_MBlog posts, structured contentSolid, controllable
8 GBMistral 7B InstructQ4_K_MQuick drafts, brainstormingFast, serviceable
12 GBMistral Nemo 12B + mergesQ4_K_MCharacter-driven fictionThe 12B fiction champ (see below)
12 GBGemma 4 12BQ4_K_MBlog posts, multimodal, editingClean, current-gen
16 GBQwen3-14BQ4_K_MBalanced creative + factualBest value mid-range
16 GBGemma 4 12BQ5_K_MArticles, SEO content, editingVery good, 256K context
24 GBQwen3 32BQ4_K_MFiction, long-form, editingExcellent โ€” the sweet spot
24 GBGemma 4 31BQ4_K_MMultimodal, general creativeStrong generalist, 256K ctx
24 GBMistral Small 24BQ4_K_MNonfiction, editing, rewritingReliable, structured
48 GB+Llama 3.3 70BQ4_K_MBest all-round local proseThe 2026 consensus top pick
48 GB+Midnight Miqu 70B v1.5Q4_K_MLiterary fiction, prose voiceBeloved classic, now aging

The jump that matters: 32B is where writing quality shifts from “useful assistant” to “genuinely good collaborator.” If you’re serious about using local AI for writing, a 24GB GPU running Qwen3 32B is the target โ€” and if you can reach 48GB, Llama 3.3 70B is the step up.

โ†’ Check what fits your hardware with our Planning Tool.


Best for Fiction & Creative Writing

Fiction is the hardest test for a language model. It needs to maintain character voice, pace a scene, build tension, and produce prose that doesn’t sound like a corporate memo.

Top picks:

TierModelWhy
Best overall (if hardware allows)Llama 3.3 70BThe 2026 consensus best all-round local prose. Strong voice consistency, takes direction, handles dark themes when framed as fiction.
Best prose voice (the classic)Midnight Miqu 70B v1.5Long the community’s darling for pure prose “feel.” Still cited, but it’s a Miqu/Llama-2-era merge โ€” beloved and aging. A 103B version also exists.
Best on 24GBQwen3 32BStrong coherence, follows style instructions, good at sustained narrative. Nearly Llama-70B prose on far less VRAM.
Best on 24GB (multimodal/general)Gemma 4 31BGoogle’s current-gen dense model, 256K context. A capable generalist for creative work, if not a dedicated prose specialist.
Best on 12GBMistral Nemo 12B + creative mergesNemo is still the 12B fiction base. Current community merges โ€” Mag-Mell 12B, Magnum v2.5, Celeste, MN-WORDSTORM โ€” tune it for emotional range and character consistency.
Best on 8GBQwen3-8BCoherent long-form for its size, maintains character consistency. The pragmatic 8B fiction pick now.

What to expect by size:

  • 7-8B: Generates readable prose but tends to rush scenes, repeat phrases, and lose track of story details after a few thousand tokens.
  • 14B: Noticeably better coherence. ~30% quality improvement over 8B for long-form text per community testing. Can maintain a scene but may struggle with complex multi-character interactions.
  • 32B: Where fiction gets genuinely good. Models understand subtext, can maintain narrative threads across chapters, and produce prose with actual stylistic variety.
  • 70B: The premium tier. Natural pacing, subtlety, sustained coherence. Llama 3.3 70B is the modern all-round pick here; the beloved prose merges (Midnight Miqu, Euryale) still have devotees but are a generation-plus old now.

Best for Blog Posts & Articles

Blog writing is more structured than fiction โ€” you need clear sections, factual tone, and consistent formatting. The good news: smaller models handle this well because the structure does a lot of the heavy lifting.

Top picks:

VRAMModelWhy
8 GBLlama 3.1 8BClean, controllable prose. Good at following outline structures.
16 GBQwen3-14BDetailed, contextually complete answers. Strong instruction following.
24 GBQwen3 32BBest local option for researchy, detailed articles.
24 GBMistral Small 24BReliable, well-structured nonfiction. Fast.

Tips for blog writing with local models:

  • Provide an outline in the prompt โ€” models follow structure much better than they invent it
  • Generate section by section, not the entire article at once
  • Use a system prompt that specifies tone and audience: “Write in a direct, practical tone for technically literate readers. No filler phrases.”

Best for Editing & Rewriting

Editing is harder than generating. The model needs to understand your intent, preserve what’s good, and improve what isn’t โ€” without rewriting everything in its own voice.

Top picks:

VRAMModelWhy
12 GBPhi-4 (14B)Excels at text tasks โ€” rewriting, summarization, rephrasing. ~9GB at Q4, so it wants 12GB, not 8.
16 GBQwen3-14BStrong instruction following. Does what you ask without going rogue.
24 GBMistral Small 24BGood at targeted edits. Fast iteration.
24 GBQwen3 32BBest at complex editing tasks that require understanding context.

Key: For editing, instruction-following matters more than raw creativity. You want a model that can execute “rewrite this paragraph to be more concise while keeping the technical details” without deciding to restructure your entire article. Qwen models excel at this.


Best for Brainstorming & Outlining

Speed matters more than quality here. You want fast idea generation, not polished prose.

Any 7-8B model works well for brainstorming. Run it at Q4_K_M on 8GB VRAM and you’ll get 30-40+ tok/s โ€” fast enough for real-time conversation.

Good picks: Qwen3-8B, Llama 3.1 8B, Mistral 7B Instruct. Don’t waste 24GB of VRAM on brainstorming.


The Censorship Problem

You’re writing a thriller. A character picks up a knife. The model refuses to continue because “violence.”

This is the biggest frustration with local writing models. Default instruct models have safety filters that trigger on violence, romance, dark themes, morally complex characters, and sometimes even mild conflict. For fiction writing, this is crippling.

The Solution: Abliterated Models

Abliteration is the current standard for uncensored local models. Instead of retraining on “edgy” data (which degrades model intelligence), abliteration surgically removes the refusal mechanism from the model’s weights. The result: same intelligence, no safety refusals. The tradeoff is that an abliterated model can be slightly worse on some non-creative tasks โ€” so keep the filtered version for factual work and reach for the abliterated one when a scene needs to go somewhere the default won’t.

This guide is about which models write well. The companion question โ€” which models won’t refuse โ€” has its own guide, and it’s the one to use for tier-by-tier picks with the actual HuggingFace repos: Best Uncensored Local LLMs. As of its last update the current center of gravity is Qwen 3.6 abliterated builds, the Gemma 4 “Heretic” variants, and Dolphin 3.0 for the smaller tiers โ€” with the how-to on abliteration vs Heretic vs dataset-filtering. No point duplicating it here; if you need uncensored, start there.


System Prompts That Actually Help

The right system prompt makes a measurable difference in writing quality. Here are three that work:

For fiction writing:

You are an experienced literary fiction author. Write vivid, emotionally
engaging prose with natural dialogue. Show, don't tell. Focus on sensory
details and character psychology. Avoid these words: tapestry, delve,
testament, beacon, journey, realm. Avoid single-sentence paragraphs.
Do not summarize emotions โ€” show them through action and dialogue.
Do not rush to resolution. Build scenes gradually.

For blog/article writing:

You are a technical writer who explains complex topics clearly. Write
in a direct, practical tone. Lead with the useful information, not
background context. Use short paragraphs. No filler phrases like
"it's worth noting" or "in today's landscape." Be specific โ€” use
numbers, examples, and concrete details instead of vague claims.

For editing:

You are a careful editor. When asked to edit text, preserve the
author's voice and intent. Only change what is specifically requested.
Do not add new content unless asked. Do not restructure unless asked.
Explain each change you make and why.

Critical tip: Repeat your most important instructions. Models deprioritize instructions over long conversations. If “show, don’t tell” matters, say it in the system prompt and again in your scene prompt.


Context Length and Long-Form Writing

Models advertise 128K context windows, but writing quality degrades well before that. Research consistently shows performance drops start at 8K-16K tokens, even when models can technically handle more.

Practical limits for writing:

ContextPagesWhat WorksWhat Breaks
2-4K tokens3-6 pagesSingle scenes, dialogueโ€”
8-16K tokens12-24 pagesChapters, short storiesCharacter details start drifting
16-32K tokens24-49 pagesMulti-chapter with summariesTone inconsistency, repetition
32K+ tokens49+ pagesPossible but quality suffersNarrative coherence degrades

The practical approach: Write chapter by chapter. Keep a running summary of characters, plot points, and tone in the system prompt. Feed the summary + current chapter into context rather than trying to keep the entire manuscript loaded. Most professional AI-assisted writers cap generation at 800-1200 words per turn for best quality.

Set num_ctx appropriately โ€” don’t rely on defaults. On constrained VRAM, KV cache quantization (Q8) can double your usable context with minimal quality loss.


Hardware Quick Reference

Writing GoalMinimum VRAMRecommended GPUModel to Run
Brainstorming, outlining8 GBRTX 3060 8GB, RX 6600Any 7-8B
Blog posts, articles12-16 GBRTX 3060 12GBQwen3-14B
Serious fiction24 GBRTX 3090 (used, ~$1,200)Qwen3 32B
Best local prose40-48 GBDual 3090 or Mac StudioLlama 3.3 70B
CPU-only option16-32 GB RAMโ€”7-14B at Q4

The Bottom Line

Local models are genuinely useful for writing now โ€” not just as gimmicks, but as real tools. The key is matching the right model to the right task:

  • Brainstorming: Any 7-8B model. Speed over quality.
  • Blog posts: Qwen3-14B on 16GB VRAM. Structured, reliable.
  • Fiction: Qwen3 32B on 24GB VRAM. The sweet spot where prose gets good.
  • Best prose: Llama 3.3 70B if you have the hardware. The 2026 all-round pick; the older prose merges are a matter of taste now.
  • Uncensored: See the uncensored models guide โ€” Qwen 3.6 abliterated, Gemma 4 Heretic, Dolphin 3.0.

Don’t fight safety filters with clever prompting โ€” switch to an abliterated model. Don’t try to generate entire novels in one shot โ€” work chapter by chapter. And if your writing feels generic, go bigger. The 14B-to-32B jump is where local AI stops sounding like an AI.