Best local LLMs for coding
Models tuned for code generation, completion and refactoring — the ones worth pointing your editor or agent at.
Best pick by hardware
The strongest coding model that runs well on each tier (at 8K context).
| If you have | Best coding model | Quant | Needs | Speed |
|---|---|---|---|---|
| RTX 4060 Ti 8 GB | Qwen2.5-Coder 7B | Q6 | 7.2 GB | ~33 tok/s |
| RTX 3060 12 GB | Qwen3 14B | Q4 | 11 GB | ~29 tok/s |
| RTX 4060 Ti 16 GB | Mistral Small 3 24B | Q3 | 13 GB | ~18 tok/s |
| RTX 4090 | Qwen3 32B | Q4 | 22 GB | ~37 tok/s |
| RTX A6000 | Qwen2.5 72B | Q4 | 46 GB | ~13 tok/s |
| Mac · M4 Max, 64 GB | Qwen2.5 72B | Q4 | 46 GB | ~9.0 tok/s |
All coding models, ranked by size
- 1. DeepSeek-R1 671B-A37B (MoE) 671B · MIT
The closest open weights to frontier reasoning. ~400 GB even at Q4 — a server or a maxed multi-GPU rig only.
- 2. Qwen3 235B-A22B (MoE) 235B · Apache 2.0
Frontier-class open weights. ~140 GB at Q4 — a 192 GB Mac Studio or a multi-GPU server. Fast for its size.
- 3. Qwen2.5 72B 72.7B · Qwen
Among the strongest open dense models. Same memory class as Llama 70B; pick by benchmark for your task.
- 4. Llama 3.3 70B 70.6B · Llama 3.3 Community
The classic 70B target. Needs ~40 GB at Q4 — dual 24 GB cards, a 48 GB card, or a 64 GB+ Mac.
- 5. DeepSeek-R1 Distill Llama 70B 70.6B · Llama 3.3 Community
The best open reasoning you can run locally short of the full R1 — if you have ~40 GB of memory.
- 6. Qwen3 32B 32.8B · Apache 2.0
The flagship dense model for a 24 GB card. Thinking mode + tool use; the best single-GPU all-rounder of 2025.
- 7. Qwen2.5 32B 32.8B · Apache 2.0
The dependable 32B with the deepest pool of fine-tunes. A safe 24 GB-card choice.
- 8. Qwen2.5-Coder 32B 32.8B · Apache 2.0
The best local coding model that fits one 24 GB card. Trades blows with hosted models on many coding tasks.
- 9. DeepSeek-R1 Distill Qwen 32B 32.8B · MIT
The strongest reasoning model that fits a single 24 GB card. The go-to local "thinker".
- 10. Mistral Small 3 24B 23.6B · Apache 2.0
Near-70B feel at 24B, Apache-licensed. A superb fit for a 24 GB card or a 32 GB Mac.
- 11. Qwen3 14B 14.8B · Apache 2.0
The best all-round model that fits a single 12 GB card at Q4. Thinking mode, tool use, long context.
- 12. Qwen2.5 14B 14.8B · Apache 2.0
The proven 14B before Qwen3. Still excellent and has more fine-tunes available today.
- 13. DeepSeek-R1 Distill Qwen 14B 14.8B · MIT
R1’s reasoning distilled into a 14B that fits a 12 GB card. Pick this for math and step-by-step problems.
- 14. Qwen2.5-Coder 7B 7.6B · Apache 2.0
The best coding model that fits 8 GB. Strong fill-in-the-middle for editor autocomplete.
FAQ
What is the best local LLM for coding right now?
DeepSeek-R1 671B-A37B (MoE) is our top pick for coding: The closest open weights to frontier reasoning. ~400 GB even at Q4 — a server or a maxed multi-GPU rig only.
What can I run for coding on a 24 GB GPU?
On a 24 GB card like the RTX 4090, Qwen3 32B (Q4) is the strongest coding model that fits well.
Rankings use parameter count as a capability proxy and our computed fit — see methodology. Pick by benchmark for your exact task.