Best local LLMs for general chat & assistance

All-rounders for everyday questions, drafting and summarising — the models to reach for when you just want a capable local assistant.

Best pick by hardware

The strongest general chat & assistance model that runs well on each tier (at 8K context).

If you haveBest general chat & assistance modelQuantNeedsSpeed
RTX 4060 Ti 8 GB Mistral Nemo 12B Q2 6.9 GB ~41 tok/s
RTX 3060 12 GB Qwen3 14B Q4 11 GB ~29 tok/s
RTX 4060 Ti 16 GB Qwen3 30B-A3B (MoE) Q2 14 GB ~98 tok/s
RTX 4090 Mixtral 8x7B (MoE) Q2 21 GB ~87 tok/s
RTX A6000 Llama 4 Scout (MoE) Q2 46 GB ~50 tok/s
Mac · M4 Max, 64 GB Qwen2.5 72B Q4 46 GB ~9.0 tok/s

All general chat & assistance models, ranked by size

  1. The closest open weights to frontier reasoning. ~400 GB even at Q4 — a server or a maxed multi-GPU rig only.

  2. 2. Qwen3 235B-A22B (MoE) 235B · Apache 2.0

    Frontier-class open weights. ~140 GB at Q4 — a 192 GB Mac Studio or a multi-GPU server. Fast for its size.

  3. 3. Llama 4 Scout (MoE) 109B · Llama 4 Community

    A 109B MoE with only 17B active and a huge context. Wants ~60 GB+ — a 96–128 GB Mac or multi-GPU box.

  4. 4. Qwen2.5 72B 72.7B · Qwen

    Among the strongest open dense models. Same memory class as Llama 70B; pick by benchmark for your task.

  5. 5. Llama 3.3 70B 70.6B · Llama 3.3 Community

    The classic 70B target. Needs ~40 GB at Q4 — dual 24 GB cards, a 48 GB card, or a 64 GB+ Mac.

  6. 6. Mixtral 8x7B (MoE) 46.7B · Apache 2.0

    The original open MoE. Needs ~28 GB at Q4 but runs fast for its quality. Aging, but still capable.

  7. 7. Qwen3 32B 32.8B · Apache 2.0

    The flagship dense model for a 24 GB card. Thinking mode + tool use; the best single-GPU all-rounder of 2025.

  8. 8. Qwen2.5 32B 32.8B · Apache 2.0

    The dependable 32B with the deepest pool of fine-tunes. A safe 24 GB-card choice.

  9. 9. Qwen3 30B-A3B (MoE) 30.5B · Apache 2.0

    Holds 30B in memory but only computes 3B per token — 14B-class quality at near-8B speed. Loves Macs.

  10. 10. Gemma 3 27B 27.4B · Gemma

    The best open multimodal model you can run on one 24 GB card at Q4. Excellent writing, 128K context.

  11. 11. Gemma 2 27B 27.2B · Gemma

    Still a strong writer, but the 8K context and lack of vision make Gemma 3 27B the better pick today.

  12. 12. Mistral Small 3 24B 23.6B · Apache 2.0

    Near-70B feel at 24B, Apache-licensed. A superb fit for a 24 GB card or a 32 GB Mac.

  13. 13. Qwen3 14B 14.8B · Apache 2.0

    The best all-round model that fits a single 12 GB card at Q4. Thinking mode, tool use, long context.

  14. 14. Qwen2.5 14B 14.8B · Apache 2.0

    The proven 14B before Qwen3. Still excellent and has more fine-tunes available today.

  15. 15. Phi-4 14B 14.7B · MIT

    Microsoft’s small reasoner — outpunches its size on math and logic. MIT-licensed. Modest 16K context.

  16. 16. Mistral Nemo 12B 12.2B · Apache 2.0

    A 12B with a real 128K context and a permissive licence — a sweet spot for a 12 GB card.

  17. 17. Gemma 3 12B 12.2B · Gemma

    Multimodal, strong writing, 128K context. The 12 GB-card pick when you want vision and long context.

  18. 18. Gemma 2 9B 9.2B · Gemma

    Unusually good prose for its size. Short 8K context is the catch — great for chat, weak for long documents.

  19. 19. Qwen3 8B 8.2B · Apache 2.0

    A 2025 8B with a toggleable thinking mode — currently the strongest all-rounder in the 8 GB class.

  20. 20. Llama 3.1 8B 8.0B · Llama 3.1 Community

    The default 8B everyone benchmarks against. Fits a 6–8 GB card at Q4 and just works for general chat.

  21. 21. Mistral 7B v0.3 7.3B · Apache 2.0

    The old reliable. Apache-licensed, fast, uncensored fine-tunes everywhere — still a fine 8 GB workhorse.

  22. 22. Gemma 3 4B 4.3B · Gemma

    Small and multimodal — it can read images. A good edge pick when you need vision, not just text.

  23. 23. Qwen3 4B 4B · Apache 2.0

    Punches far above 4B thanks to a thinking mode. The best tiny model for reasoning on edge hardware.

  24. 24. Llama 3.2 3B 3.2B · Llama 3.2 Community

    Pick this if you want a snappy assistant on a laptop or 8 GB card and can live with the odd mistake.

  25. 25. Llama 3.2 1B 1.2B · Llama 3.2 Community

    The smallest genuinely useful Llama. Runs on almost anything — phones, Raspberry Pi, CPU — for classification and simple chat.

FAQ

What is the best local LLM for general chat & assistance right now?

DeepSeek-R1 671B-A37B (MoE) is our top pick for general chat & assistance: The closest open weights to frontier reasoning. ~400 GB even at Q4 — a server or a maxed multi-GPU rig only.

What can I run for general chat & assistance on a 24 GB GPU?

On a 24 GB card like the RTX 4090, Mixtral 8x7B (MoE) (Q2) is the strongest general chat & assistance model that fits well.

Rankings use parameter count as a capability proxy and our computed fit — see methodology. Pick by benchmark for your exact task.