Best local LLMs for tool use & agents

Models that reliably follow instructions and call tools — the ones to build local agents and automations on.

Best pick by hardware

The strongest tool use & agents model that runs well on each tier (at 8K context).

If you haveBest tool use & agents modelQuantNeedsSpeed
RTX 4060 Ti 8 GB Qwen3 8B Q4 6.7 GB ~42 tok/s
RTX 3060 12 GB Qwen3 14B Q4 11 GB ~29 tok/s
RTX 4060 Ti 16 GB Qwen3 30B-A3B (MoE) Q2 14 GB ~98 tok/s
RTX 4090 Qwen3 32B Q4 22 GB ~37 tok/s
RTX A6000 Llama 4 Scout (MoE) Q2 46 GB ~50 tok/s
Mac · M4 Max, 64 GB Llama 3.3 70B Q4 45 GB ~9.2 tok/s

All tool use & agents models, ranked by size

  1. 1. Qwen3 235B-A22B (MoE) 235B · Apache 2.0

    Frontier-class open weights. ~140 GB at Q4 — a 192 GB Mac Studio or a multi-GPU server. Fast for its size.

  2. 2. Llama 4 Scout (MoE) 109B · Llama 4 Community

    A 109B MoE with only 17B active and a huge context. Wants ~60 GB+ — a 96–128 GB Mac or multi-GPU box.

  3. 3. Llama 3.3 70B 70.6B · Llama 3.3 Community

    The classic 70B target. Needs ~40 GB at Q4 — dual 24 GB cards, a 48 GB card, or a 64 GB+ Mac.

  4. 4. Qwen3 32B 32.8B · Apache 2.0

    The flagship dense model for a 24 GB card. Thinking mode + tool use; the best single-GPU all-rounder of 2025.

  5. 5. Qwen2.5-Coder 32B 32.8B · Apache 2.0

    The best local coding model that fits one 24 GB card. Trades blows with hosted models on many coding tasks.

  6. 6. Qwen3 30B-A3B (MoE) 30.5B · Apache 2.0

    Holds 30B in memory but only computes 3B per token — 14B-class quality at near-8B speed. Loves Macs.

  7. 7. Mistral Small 3 24B 23.6B · Apache 2.0

    Near-70B feel at 24B, Apache-licensed. A superb fit for a 24 GB card or a 32 GB Mac.

  8. 8. Qwen3 14B 14.8B · Apache 2.0

    The best all-round model that fits a single 12 GB card at Q4. Thinking mode, tool use, long context.

  9. 9. Qwen3 8B 8.2B · Apache 2.0

    A 2025 8B with a toggleable thinking mode — currently the strongest all-rounder in the 8 GB class.

  10. 10. Llama 3.1 8B 8.0B · Llama 3.1 Community

    The default 8B everyone benchmarks against. Fits a 6–8 GB card at Q4 and just works for general chat.

  11. 11. Qwen2.5-Coder 7B 7.6B · Apache 2.0

    The best coding model that fits 8 GB. Strong fill-in-the-middle for editor autocomplete.

  12. 12. Llama 3.2 3B 3.2B · Llama 3.2 Community

    Pick this if you want a snappy assistant on a laptop or 8 GB card and can live with the odd mistake.

FAQ

What is the best local LLM for tool use & agents right now?

Qwen3 235B-A22B (MoE) is our top pick for tool use & agents: Frontier-class open weights. ~140 GB at Q4 — a 192 GB Mac Studio or a multi-GPU server. Fast for its size.

What can I run for tool use & agents on a 24 GB GPU?

On a 24 GB card like the RTX 4090, Qwen3 32B (Q4) is the strongest tool use & agents model that fits well.

Rankings use parameter count as a capability proxy and our computed fit — see methodology. Pick by benchmark for your exact task.