Best local LLMs for vision (image input)
Multimodal models that can read images, screenshots and documents alongside text.
Best pick by hardware
The strongest vision (image input) model that runs well on each tier (at 8K context).
| If you have | Best vision (image input) model | Quant | Needs | Speed |
|---|---|---|---|---|
| RTX 4060 Ti 8 GB | Gemma 3 4B | Q8 | 6.2 GB | ~45 tok/s |
| RTX 3060 12 GB | Gemma 3 12B | Q4 | 11 GB | ~35 tok/s |
| RTX 4060 Ti 16 GB | Gemma 3 12B | Q6 | 13 GB | ~21 tok/s |
| RTX 4090 | Gemma 3 27B | Q4 | 21 GB | ~44 tok/s |
| RTX A6000 | Llama 4 Scout (MoE) | Q2 | 46 GB | ~50 tok/s |
| Mac · M4 Max, 64 GB | Gemma 3 27B | Q8 | 33 GB | ~14 tok/s |
All vision (image input) models, ranked by size
- 1. Llama 4 Scout (MoE) 109B · Llama 4 Community
A 109B MoE with only 17B active and a huge context. Wants ~60 GB+ — a 96–128 GB Mac or multi-GPU box.
- 2. Gemma 3 27B 27.4B · Gemma
The best open multimodal model you can run on one 24 GB card at Q4. Excellent writing, 128K context.
- 3. Gemma 3 12B 12.2B · Gemma
Multimodal, strong writing, 128K context. The 12 GB-card pick when you want vision and long context.
- 4. Gemma 3 4B 4.3B · Gemma
Small and multimodal — it can read images. A good edge pick when you need vision, not just text.
FAQ
What is the best local LLM for vision (image input) right now?
Llama 4 Scout (MoE) is our top pick for vision (image input): A 109B MoE with only 17B active and a huge context. Wants ~60 GB+ — a 96–128 GB Mac or multi-GPU box.
What can I run for vision (image input) on a 24 GB GPU?
On a 24 GB card like the RTX 4090, Gemma 3 27B (Q4) is the strongest vision (image input) model that fits well.
Rankings use parameter count as a capability proxy and our computed fit — see methodology. Pick by benchmark for your exact task.