Best local LLMs for general chat & assistance
All-rounders for everyday questions, drafting and summarising — the models to reach for when you just want a capable local assistant.
Best pick by hardware
The strongest general chat & assistance model that runs well on each tier (at 8K context).
| If you have | Best general chat & assistance model | Quant | Needs | Speed |
|---|---|---|---|---|
| RTX 4060 Ti 8 GB | Mistral Nemo 12B | Q2 | 6.9 GB | ~41 tok/s |
| RTX 3060 12 GB | Qwen3 14B | Q4 | 11 GB | ~29 tok/s |
| RTX 4060 Ti 16 GB | Qwen3 30B-A3B (MoE) | Q2 | 14 GB | ~98 tok/s |
| RTX 4090 | Mixtral 8x7B (MoE) | Q2 | 21 GB | ~87 tok/s |
| RTX A6000 | Llama 4 Scout (MoE) | Q2 | 46 GB | ~50 tok/s |
| Mac · M4 Max, 64 GB | Qwen2.5 72B | Q4 | 46 GB | ~9.0 tok/s |
All general chat & assistance models, ranked by size
- 1. DeepSeek-R1 671B-A37B (MoE) 671B · MIT
The closest open weights to frontier reasoning. ~400 GB even at Q4 — a server or a maxed multi-GPU rig only.
- 2. Qwen3 235B-A22B (MoE) 235B · Apache 2.0
Frontier-class open weights. ~140 GB at Q4 — a 192 GB Mac Studio or a multi-GPU server. Fast for its size.
- 3. Llama 4 Scout (MoE) 109B · Llama 4 Community
A 109B MoE with only 17B active and a huge context. Wants ~60 GB+ — a 96–128 GB Mac or multi-GPU box.
- 4. Qwen2.5 72B 72.7B · Qwen
Among the strongest open dense models. Same memory class as Llama 70B; pick by benchmark for your task.
- 5. Llama 3.3 70B 70.6B · Llama 3.3 Community
The classic 70B target. Needs ~40 GB at Q4 — dual 24 GB cards, a 48 GB card, or a 64 GB+ Mac.
- 6. Mixtral 8x7B (MoE) 46.7B · Apache 2.0
The original open MoE. Needs ~28 GB at Q4 but runs fast for its quality. Aging, but still capable.
- 7. Qwen3 32B 32.8B · Apache 2.0
The flagship dense model for a 24 GB card. Thinking mode + tool use; the best single-GPU all-rounder of 2025.
- 8. Qwen2.5 32B 32.8B · Apache 2.0
The dependable 32B with the deepest pool of fine-tunes. A safe 24 GB-card choice.
- 9. Qwen3 30B-A3B (MoE) 30.5B · Apache 2.0
Holds 30B in memory but only computes 3B per token — 14B-class quality at near-8B speed. Loves Macs.
- 10. Gemma 3 27B 27.4B · Gemma
The best open multimodal model you can run on one 24 GB card at Q4. Excellent writing, 128K context.
- 11. Gemma 2 27B 27.2B · Gemma
Still a strong writer, but the 8K context and lack of vision make Gemma 3 27B the better pick today.
- 12. Mistral Small 3 24B 23.6B · Apache 2.0
Near-70B feel at 24B, Apache-licensed. A superb fit for a 24 GB card or a 32 GB Mac.
- 13. Qwen3 14B 14.8B · Apache 2.0
The best all-round model that fits a single 12 GB card at Q4. Thinking mode, tool use, long context.
- 14. Qwen2.5 14B 14.8B · Apache 2.0
The proven 14B before Qwen3. Still excellent and has more fine-tunes available today.
- 15. Phi-4 14B 14.7B · MIT
Microsoft’s small reasoner — outpunches its size on math and logic. MIT-licensed. Modest 16K context.
- 16. Mistral Nemo 12B 12.2B · Apache 2.0
A 12B with a real 128K context and a permissive licence — a sweet spot for a 12 GB card.
- 17. Gemma 3 12B 12.2B · Gemma
Multimodal, strong writing, 128K context. The 12 GB-card pick when you want vision and long context.
- 18. Gemma 2 9B 9.2B · Gemma
Unusually good prose for its size. Short 8K context is the catch — great for chat, weak for long documents.
- 19. Qwen3 8B 8.2B · Apache 2.0
A 2025 8B with a toggleable thinking mode — currently the strongest all-rounder in the 8 GB class.
- 20. Llama 3.1 8B 8.0B · Llama 3.1 Community
The default 8B everyone benchmarks against. Fits a 6–8 GB card at Q4 and just works for general chat.
- 21. Mistral 7B v0.3 7.3B · Apache 2.0
The old reliable. Apache-licensed, fast, uncensored fine-tunes everywhere — still a fine 8 GB workhorse.
- 22. Gemma 3 4B 4.3B · Gemma
Small and multimodal — it can read images. A good edge pick when you need vision, not just text.
- 23. Qwen3 4B 4B · Apache 2.0
Punches far above 4B thanks to a thinking mode. The best tiny model for reasoning on edge hardware.
- 24. Llama 3.2 3B 3.2B · Llama 3.2 Community
Pick this if you want a snappy assistant on a laptop or 8 GB card and can live with the odd mistake.
- 25. Llama 3.2 1B 1.2B · Llama 3.2 Community
The smallest genuinely useful Llama. Runs on almost anything — phones, Raspberry Pi, CPU — for classification and simple chat.
FAQ
What is the best local LLM for general chat & assistance right now?
DeepSeek-R1 671B-A37B (MoE) is our top pick for general chat & assistance: The closest open weights to frontier reasoning. ~400 GB even at Q4 — a server or a maxed multi-GPU rig only.
What can I run for general chat & assistance on a 24 GB GPU?
On a 24 GB card like the RTX 4090, Mixtral 8x7B (MoE) (Q2) is the strongest general chat & assistance model that fits well.
Rankings use parameter count as a capability proxy and our computed fit — see methodology. Pick by benchmark for your exact task.