Best local LLMs for reasoning & math
Reasoning ("thinking") models that work through problems step by step before answering — stronger on math, logic and multi-step tasks.
Best pick by hardware
The strongest reasoning & math model that runs well on each tier (at 8K context).
| If you have | Best reasoning & math model | Quant | Needs | Speed |
|---|---|---|---|---|
| RTX 4060 Ti 8 GB | Qwen3 8B | Q4 | 6.7 GB | ~42 tok/s |
| RTX 3060 12 GB | Qwen3 14B | Q4 | 11 GB | ~29 tok/s |
| RTX 4060 Ti 16 GB | Qwen3 30B-A3B (MoE) | Q2 | 14 GB | ~98 tok/s |
| RTX 4090 | Qwen3 32B | Q4 | 22 GB | ~37 tok/s |
| RTX A6000 | Qwen2.5 72B | Q4 | 46 GB | ~13 tok/s |
| Mac · M4 Max, 64 GB | Qwen2.5 72B | Q4 | 46 GB | ~9.0 tok/s |
All reasoning & math models, ranked by size
- 1. DeepSeek-R1 671B-A37B (MoE) 671B · MIT
The closest open weights to frontier reasoning. ~400 GB even at Q4 — a server or a maxed multi-GPU rig only.
- 2. Qwen3 235B-A22B (MoE) 235B · Apache 2.0
Frontier-class open weights. ~140 GB at Q4 — a 192 GB Mac Studio or a multi-GPU server. Fast for its size.
- 3. Qwen2.5 72B 72.7B · Qwen
Among the strongest open dense models. Same memory class as Llama 70B; pick by benchmark for your task.
- 4. DeepSeek-R1 Distill Llama 70B 70.6B · Llama 3.3 Community
The best open reasoning you can run locally short of the full R1 — if you have ~40 GB of memory.
- 5. Qwen3 32B 32.8B · Apache 2.0
The flagship dense model for a 24 GB card. Thinking mode + tool use; the best single-GPU all-rounder of 2025.
- 6. DeepSeek-R1 Distill Qwen 32B 32.8B · MIT
The strongest reasoning model that fits a single 24 GB card. The go-to local "thinker".
- 7. Qwen3 30B-A3B (MoE) 30.5B · Apache 2.0
Holds 30B in memory but only computes 3B per token — 14B-class quality at near-8B speed. Loves Macs.
- 8. Qwen3 14B 14.8B · Apache 2.0
The best all-round model that fits a single 12 GB card at Q4. Thinking mode, tool use, long context.
- 9. DeepSeek-R1 Distill Qwen 14B 14.8B · MIT
R1’s reasoning distilled into a 14B that fits a 12 GB card. Pick this for math and step-by-step problems.
- 10. Phi-4 14B 14.7B · MIT
Microsoft’s small reasoner — outpunches its size on math and logic. MIT-licensed. Modest 16K context.
- 11. Qwen3 8B 8.2B · Apache 2.0
A 2025 8B with a toggleable thinking mode — currently the strongest all-rounder in the 8 GB class.
- 12. Qwen3 4B 4B · Apache 2.0
Punches far above 4B thanks to a thinking mode. The best tiny model for reasoning on edge hardware.
FAQ
What is the best local LLM for reasoning & math right now?
DeepSeek-R1 671B-A37B (MoE) is our top pick for reasoning & math: The closest open weights to frontier reasoning. ~400 GB even at Q4 — a server or a maxed multi-GPU rig only.
What can I run for reasoning & math on a 24 GB GPU?
On a 24 GB card like the RTX 4090, Qwen3 32B (Q4) is the strongest reasoning & math model that fits well.
Rankings use parameter count as a capability proxy and our computed fit — see methodology. Pick by benchmark for your exact task.