Best local LLMs for low-end & edge hardware
Small models that stay useful on laptops, mini-PCs and 8 GB cards — when memory is the hard constraint.
Best pick by hardware
The strongest low-end & edge hardware model that runs well on each tier (at 8K context).
| If you have | Best low-end & edge hardware model | Quant | Needs | Speed |
|---|---|---|---|---|
| RTX 4060 Ti 8 GB | Gemma 3 4B | Q8 | 6.2 GB | ~45 tok/s |
| RTX 3060 12 GB | Gemma 3 4B | FP16 | 10 GB | ~30 tok/s |
| RTX 4060 Ti 16 GB | Gemma 3 4B | FP16 | 10 GB | ~24 tok/s |
| RTX 4090 | Gemma 3 4B | FP16 | 10 GB | ~84 tok/s |
| RTX A6000 | Gemma 3 4B | FP16 | 10 GB | ~64 tok/s |
| Mac · M4 Max, 64 GB | Gemma 3 4B | FP16 | 10 GB | ~46 tok/s |
All low-end & edge hardware models, ranked by size
- 1. Gemma 3 4B 4.3B · Gemma
Small and multimodal — it can read images. A good edge pick when you need vision, not just text.
- 2. Qwen3 4B 4B · Apache 2.0
Punches far above 4B thanks to a thinking mode. The best tiny model for reasoning on edge hardware.
- 3. Llama 3.2 3B 3.2B · Llama 3.2 Community
Pick this if you want a snappy assistant on a laptop or 8 GB card and can live with the odd mistake.
- 4. Llama 3.2 1B 1.2B · Llama 3.2 Community
The smallest genuinely useful Llama. Runs on almost anything — phones, Raspberry Pi, CPU — for classification and simple chat.
FAQ
What is the best local LLM for low-end & edge hardware right now?
Gemma 3 4B is our top pick for low-end & edge hardware: Small and multimodal — it can read images. A good edge pick when you need vision, not just text.
What can I run for low-end & edge hardware on a 24 GB GPU?
On a 24 GB card like the RTX 4090, Gemma 3 4B (FP16) is the strongest low-end & edge hardware model that fits well.
Rankings use parameter count as a capability proxy and our computed fit — see methodology. Pick by benchmark for your exact task.