Best local LLMs for low-end & edge hardware

Small models that stay useful on laptops, mini-PCs and 8 GB cards — when memory is the hard constraint.

Best pick by hardware

The strongest low-end & edge hardware model that runs well on each tier (at 8K context).

If you haveBest low-end & edge hardware modelQuantNeedsSpeed
RTX 4060 Ti 8 GB Gemma 3 4B Q8 6.2 GB ~45 tok/s
RTX 3060 12 GB Gemma 3 4B FP16 10 GB ~30 tok/s
RTX 4060 Ti 16 GB Gemma 3 4B FP16 10 GB ~24 tok/s
RTX 4090 Gemma 3 4B FP16 10 GB ~84 tok/s
RTX A6000 Gemma 3 4B FP16 10 GB ~64 tok/s
Mac · M4 Max, 64 GB Gemma 3 4B FP16 10 GB ~46 tok/s

All low-end & edge hardware models, ranked by size

  1. 1. Gemma 3 4B 4.3B · Gemma

    Small and multimodal — it can read images. A good edge pick when you need vision, not just text.

  2. 2. Qwen3 4B 4B · Apache 2.0

    Punches far above 4B thanks to a thinking mode. The best tiny model for reasoning on edge hardware.

  3. 3. Llama 3.2 3B 3.2B · Llama 3.2 Community

    Pick this if you want a snappy assistant on a laptop or 8 GB card and can live with the odd mistake.

  4. 4. Llama 3.2 1B 1.2B · Llama 3.2 Community

    The smallest genuinely useful Llama. Runs on almost anything — phones, Raspberry Pi, CPU — for classification and simple chat.

FAQ

What is the best local LLM for low-end & edge hardware right now?

Gemma 3 4B is our top pick for low-end & edge hardware: Small and multimodal — it can read images. A good edge pick when you need vision, not just text.

What can I run for low-end & edge hardware on a 24 GB GPU?

On a 24 GB card like the RTX 4090, Gemma 3 4B (FP16) is the strongest low-end & edge hardware model that fits well.

Rankings use parameter count as a capability proxy and our computed fit — see methodology. Pick by benchmark for your exact task.