Best local LLMs for tool use & agents
Models that reliably follow instructions and call tools — the ones to build local agents and automations on.
Best pick by hardware
The strongest tool use & agents model that runs well on each tier (at 8K context).
| If you have | Best tool use & agents model | Quant | Needs | Speed |
|---|---|---|---|---|
| RTX 4060 Ti 8 GB | Qwen3 8B | Q4 | 6.7 GB | ~42 tok/s |
| RTX 3060 12 GB | Qwen3 14B | Q4 | 11 GB | ~29 tok/s |
| RTX 4060 Ti 16 GB | Qwen3 30B-A3B (MoE) | Q2 | 14 GB | ~98 tok/s |
| RTX 4090 | Qwen3 32B | Q4 | 22 GB | ~37 tok/s |
| RTX A6000 | Llama 4 Scout (MoE) | Q2 | 46 GB | ~50 tok/s |
| Mac · M4 Max, 64 GB | Llama 3.3 70B | Q4 | 45 GB | ~9.2 tok/s |
All tool use & agents models, ranked by size
- 1. Qwen3 235B-A22B (MoE) 235B · Apache 2.0
Frontier-class open weights. ~140 GB at Q4 — a 192 GB Mac Studio or a multi-GPU server. Fast for its size.
- 2. Llama 4 Scout (MoE) 109B · Llama 4 Community
A 109B MoE with only 17B active and a huge context. Wants ~60 GB+ — a 96–128 GB Mac or multi-GPU box.
- 3. Llama 3.3 70B 70.6B · Llama 3.3 Community
The classic 70B target. Needs ~40 GB at Q4 — dual 24 GB cards, a 48 GB card, or a 64 GB+ Mac.
- 4. Qwen3 32B 32.8B · Apache 2.0
The flagship dense model for a 24 GB card. Thinking mode + tool use; the best single-GPU all-rounder of 2025.
- 5. Qwen2.5-Coder 32B 32.8B · Apache 2.0
The best local coding model that fits one 24 GB card. Trades blows with hosted models on many coding tasks.
- 6. Qwen3 30B-A3B (MoE) 30.5B · Apache 2.0
Holds 30B in memory but only computes 3B per token — 14B-class quality at near-8B speed. Loves Macs.
- 7. Mistral Small 3 24B 23.6B · Apache 2.0
Near-70B feel at 24B, Apache-licensed. A superb fit for a 24 GB card or a 32 GB Mac.
- 8. Qwen3 14B 14.8B · Apache 2.0
The best all-round model that fits a single 12 GB card at Q4. Thinking mode, tool use, long context.
- 9. Qwen3 8B 8.2B · Apache 2.0
A 2025 8B with a toggleable thinking mode — currently the strongest all-rounder in the 8 GB class.
- 10. Llama 3.1 8B 8.0B · Llama 3.1 Community
The default 8B everyone benchmarks against. Fits a 6–8 GB card at Q4 and just works for general chat.
- 11. Qwen2.5-Coder 7B 7.6B · Apache 2.0
The best coding model that fits 8 GB. Strong fill-in-the-middle for editor autocomplete.
- 12. Llama 3.2 3B 3.2B · Llama 3.2 Community
Pick this if you want a snappy assistant on a laptop or 8 GB card and can live with the odd mistake.
FAQ
What is the best local LLM for tool use & agents right now?
Qwen3 235B-A22B (MoE) is our top pick for tool use & agents: Frontier-class open weights. ~140 GB at Q4 — a 192 GB Mac Studio or a multi-GPU server. Fast for its size.
What can I run for tool use & agents on a 24 GB GPU?
On a 24 GB card like the RTX 4090, Qwen3 32B (Q4) is the strongest tool use & agents model that fits well.
Rankings use parameter count as a capability proxy and our computed fit — see methodology. Pick by benchmark for your exact task.