Best local LLMs for reasoning & math

Reasoning ("thinking") models that work through problems step by step before answering — stronger on math, logic and multi-step tasks.

Best pick by hardware

The strongest reasoning & math model that runs well on each tier (at 8K context).

If you haveBest reasoning & math modelQuantNeedsSpeed
RTX 4060 Ti 8 GB Qwen3 8B Q4 6.7 GB ~42 tok/s
RTX 3060 12 GB Qwen3 14B Q4 11 GB ~29 tok/s
RTX 4060 Ti 16 GB Qwen3 30B-A3B (MoE) Q2 14 GB ~98 tok/s
RTX 4090 Qwen3 32B Q4 22 GB ~37 tok/s
RTX A6000 Qwen2.5 72B Q4 46 GB ~13 tok/s
Mac · M4 Max, 64 GB Qwen2.5 72B Q4 46 GB ~9.0 tok/s

All reasoning & math models, ranked by size

  1. The closest open weights to frontier reasoning. ~400 GB even at Q4 — a server or a maxed multi-GPU rig only.

  2. 2. Qwen3 235B-A22B (MoE) 235B · Apache 2.0

    Frontier-class open weights. ~140 GB at Q4 — a 192 GB Mac Studio or a multi-GPU server. Fast for its size.

  3. 3. Qwen2.5 72B 72.7B · Qwen

    Among the strongest open dense models. Same memory class as Llama 70B; pick by benchmark for your task.

  4. 4. DeepSeek-R1 Distill Llama 70B 70.6B · Llama 3.3 Community

    The best open reasoning you can run locally short of the full R1 — if you have ~40 GB of memory.

  5. 5. Qwen3 32B 32.8B · Apache 2.0

    The flagship dense model for a 24 GB card. Thinking mode + tool use; the best single-GPU all-rounder of 2025.

  6. The strongest reasoning model that fits a single 24 GB card. The go-to local "thinker".

  7. 7. Qwen3 30B-A3B (MoE) 30.5B · Apache 2.0

    Holds 30B in memory but only computes 3B per token — 14B-class quality at near-8B speed. Loves Macs.

  8. 8. Qwen3 14B 14.8B · Apache 2.0

    The best all-round model that fits a single 12 GB card at Q4. Thinking mode, tool use, long context.

  9. R1’s reasoning distilled into a 14B that fits a 12 GB card. Pick this for math and step-by-step problems.

  10. 10. Phi-4 14B 14.7B · MIT

    Microsoft’s small reasoner — outpunches its size on math and logic. MIT-licensed. Modest 16K context.

  11. 11. Qwen3 8B 8.2B · Apache 2.0

    A 2025 8B with a toggleable thinking mode — currently the strongest all-rounder in the 8 GB class.

  12. 12. Qwen3 4B 4B · Apache 2.0

    Punches far above 4B thanks to a thinking mode. The best tiny model for reasoning on edge hardware.

FAQ

What is the best local LLM for reasoning & math right now?

DeepSeek-R1 671B-A37B (MoE) is our top pick for reasoning & math: The closest open weights to frontier reasoning. ~400 GB even at Q4 — a server or a maxed multi-GPU rig only.

What can I run for reasoning & math on a 24 GB GPU?

On a 24 GB card like the RTX 4090, Qwen3 32B (Q4) is the strongest reasoning & math model that fits well.

Rankings use parameter count as a capability proxy and our computed fit — see methodology. Pick by benchmark for your exact task.