Best local LLMs for writing
Models with the most natural prose for long-form writing, rewriting and tone control.
Best pick by hardware
The strongest writing model that runs well on each tier (at 8K context).
| If you have | Best writing model | Quant | Needs | Speed |
|---|---|---|---|---|
| RTX 4060 Ti 8 GB | Mistral Nemo 12B | Q2 | 6.9 GB | ~41 tok/s |
| RTX 3060 12 GB | Qwen2.5 14B | Q4 | 11 GB | ~29 tok/s |
| RTX 4060 Ti 16 GB | Gemma 2 27B | Q2 | 15 GB | ~18 tok/s |
| RTX 4090 | Qwen2.5 32B | Q4 | 22 GB | ~37 tok/s |
| RTX A6000 | Llama 3.3 70B | Q4 | 45 GB | ~13 tok/s |
| Mac · M4 Max, 64 GB | Llama 3.3 70B | Q4 | 45 GB | ~9.2 tok/s |
All writing models, ranked by size
- 1. Llama 3.3 70B 70.6B · Llama 3.3 Community
The classic 70B target. Needs ~40 GB at Q4 — dual 24 GB cards, a 48 GB card, or a 64 GB+ Mac.
- 2. Qwen2.5 32B 32.8B · Apache 2.0
The dependable 32B with the deepest pool of fine-tunes. A safe 24 GB-card choice.
- 3. Gemma 3 27B 27.4B · Gemma
The best open multimodal model you can run on one 24 GB card at Q4. Excellent writing, 128K context.
- 4. Gemma 2 27B 27.2B · Gemma
Still a strong writer, but the 8K context and lack of vision make Gemma 3 27B the better pick today.
- 5. Qwen2.5 14B 14.8B · Apache 2.0
The proven 14B before Qwen3. Still excellent and has more fine-tunes available today.
- 6. Mistral Nemo 12B 12.2B · Apache 2.0
A 12B with a real 128K context and a permissive licence — a sweet spot for a 12 GB card.
- 7. Gemma 3 12B 12.2B · Gemma
Multimodal, strong writing, 128K context. The 12 GB-card pick when you want vision and long context.
- 8. Gemma 2 9B 9.2B · Gemma
Unusually good prose for its size. Short 8K context is the catch — great for chat, weak for long documents.
- 9. Llama 3.1 8B 8.0B · Llama 3.1 Community
The default 8B everyone benchmarks against. Fits a 6–8 GB card at Q4 and just works for general chat.
- 10. Mistral 7B v0.3 7.3B · Apache 2.0
The old reliable. Apache-licensed, fast, uncensored fine-tunes everywhere — still a fine 8 GB workhorse.
FAQ
What is the best local LLM for writing right now?
Llama 3.3 70B is our top pick for writing: The classic 70B target. Needs ~40 GB at Q4 — dual 24 GB cards, a 48 GB card, or a 64 GB+ Mac.
What can I run for writing on a 24 GB GPU?
On a 24 GB card like the RTX 4090, Qwen2.5 32B (Q4) is the strongest writing model that fits well.
Rankings use parameter count as a capability proxy and our computed fit — see methodology. Pick by benchmark for your exact task.