Best local LLMs for writing

Models with the most natural prose for long-form writing, rewriting and tone control.

Best pick by hardware

The strongest writing model that runs well on each tier (at 8K context).

If you haveBest writing modelQuantNeedsSpeed
RTX 4060 Ti 8 GB Mistral Nemo 12B Q2 6.9 GB ~41 tok/s
RTX 3060 12 GB Qwen2.5 14B Q4 11 GB ~29 tok/s
RTX 4060 Ti 16 GB Gemma 2 27B Q2 15 GB ~18 tok/s
RTX 4090 Qwen2.5 32B Q4 22 GB ~37 tok/s
RTX A6000 Llama 3.3 70B Q4 45 GB ~13 tok/s
Mac · M4 Max, 64 GB Llama 3.3 70B Q4 45 GB ~9.2 tok/s

All writing models, ranked by size

  1. 1. Llama 3.3 70B 70.6B · Llama 3.3 Community

    The classic 70B target. Needs ~40 GB at Q4 — dual 24 GB cards, a 48 GB card, or a 64 GB+ Mac.

  2. 2. Qwen2.5 32B 32.8B · Apache 2.0

    The dependable 32B with the deepest pool of fine-tunes. A safe 24 GB-card choice.

  3. 3. Gemma 3 27B 27.4B · Gemma

    The best open multimodal model you can run on one 24 GB card at Q4. Excellent writing, 128K context.

  4. 4. Gemma 2 27B 27.2B · Gemma

    Still a strong writer, but the 8K context and lack of vision make Gemma 3 27B the better pick today.

  5. 5. Qwen2.5 14B 14.8B · Apache 2.0

    The proven 14B before Qwen3. Still excellent and has more fine-tunes available today.

  6. 6. Mistral Nemo 12B 12.2B · Apache 2.0

    A 12B with a real 128K context and a permissive licence — a sweet spot for a 12 GB card.

  7. 7. Gemma 3 12B 12.2B · Gemma

    Multimodal, strong writing, 128K context. The 12 GB-card pick when you want vision and long context.

  8. 8. Gemma 2 9B 9.2B · Gemma

    Unusually good prose for its size. Short 8K context is the catch — great for chat, weak for long documents.

  9. 9. Llama 3.1 8B 8.0B · Llama 3.1 Community

    The default 8B everyone benchmarks against. Fits a 6–8 GB card at Q4 and just works for general chat.

  10. 10. Mistral 7B v0.3 7.3B · Apache 2.0

    The old reliable. Apache-licensed, fast, uncensored fine-tunes everywhere — still a fine 8 GB workhorse.

FAQ

What is the best local LLM for writing right now?

Llama 3.3 70B is our top pick for writing: The classic 70B target. Needs ~40 GB at Q4 — dual 24 GB cards, a 48 GB card, or a 64 GB+ Mac.

What can I run for writing on a 24 GB GPU?

On a 24 GB card like the RTX 4090, Qwen2.5 32B (Q4) is the strongest writing model that fits well.

Rankings use parameter count as a capability proxy and our computed fit — see methodology. Pick by benchmark for your exact task.