Best local LLMs for coding

Models tuned for code generation, completion and refactoring — the ones worth pointing your editor or agent at.

Best pick by hardware

The strongest coding model that runs well on each tier (at 8K context).

If you haveBest coding modelQuantNeedsSpeed
RTX 4060 Ti 8 GB Qwen2.5-Coder 7B Q6 7.2 GB ~33 tok/s
RTX 3060 12 GB Qwen3 14B Q4 11 GB ~29 tok/s
RTX 4060 Ti 16 GB Mistral Small 3 24B Q3 13 GB ~18 tok/s
RTX 4090 Qwen3 32B Q4 22 GB ~37 tok/s
RTX A6000 Qwen2.5 72B Q4 46 GB ~13 tok/s
Mac · M4 Max, 64 GB Qwen2.5 72B Q4 46 GB ~9.0 tok/s

All coding models, ranked by size

  1. The closest open weights to frontier reasoning. ~400 GB even at Q4 — a server or a maxed multi-GPU rig only.

  2. 2. Qwen3 235B-A22B (MoE) 235B · Apache 2.0

    Frontier-class open weights. ~140 GB at Q4 — a 192 GB Mac Studio or a multi-GPU server. Fast for its size.

  3. 3. Qwen2.5 72B 72.7B · Qwen

    Among the strongest open dense models. Same memory class as Llama 70B; pick by benchmark for your task.

  4. 4. Llama 3.3 70B 70.6B · Llama 3.3 Community

    The classic 70B target. Needs ~40 GB at Q4 — dual 24 GB cards, a 48 GB card, or a 64 GB+ Mac.

  5. 5. DeepSeek-R1 Distill Llama 70B 70.6B · Llama 3.3 Community

    The best open reasoning you can run locally short of the full R1 — if you have ~40 GB of memory.

  6. 6. Qwen3 32B 32.8B · Apache 2.0

    The flagship dense model for a 24 GB card. Thinking mode + tool use; the best single-GPU all-rounder of 2025.

  7. 7. Qwen2.5 32B 32.8B · Apache 2.0

    The dependable 32B with the deepest pool of fine-tunes. A safe 24 GB-card choice.

  8. 8. Qwen2.5-Coder 32B 32.8B · Apache 2.0

    The best local coding model that fits one 24 GB card. Trades blows with hosted models on many coding tasks.

  9. The strongest reasoning model that fits a single 24 GB card. The go-to local "thinker".

  10. 10. Mistral Small 3 24B 23.6B · Apache 2.0

    Near-70B feel at 24B, Apache-licensed. A superb fit for a 24 GB card or a 32 GB Mac.

  11. 11. Qwen3 14B 14.8B · Apache 2.0

    The best all-round model that fits a single 12 GB card at Q4. Thinking mode, tool use, long context.

  12. 12. Qwen2.5 14B 14.8B · Apache 2.0

    The proven 14B before Qwen3. Still excellent and has more fine-tunes available today.

  13. R1’s reasoning distilled into a 14B that fits a 12 GB card. Pick this for math and step-by-step problems.

  14. 14. Qwen2.5-Coder 7B 7.6B · Apache 2.0

    The best coding model that fits 8 GB. Strong fill-in-the-middle for editor autocomplete.

FAQ

What is the best local LLM for coding right now?

DeepSeek-R1 671B-A37B (MoE) is our top pick for coding: The closest open weights to frontier reasoning. ~400 GB even at Q4 — a server or a maxed multi-GPU rig only.

What can I run for coding on a 24 GB GPU?

On a 24 GB card like the RTX 4090, Qwen3 32B (Q4) is the strongest coding model that fits well.

Rankings use parameter count as a capability proxy and our computed fit — see methodology. Pick by benchmark for your exact task.