Best local LLMs for vision (image input)

Multimodal models that can read images, screenshots and documents alongside text.

Best pick by hardware

The strongest vision (image input) model that runs well on each tier (at 8K context).

If you haveBest vision (image input) modelQuantNeedsSpeed
RTX 4060 Ti 8 GB Gemma 3 4B Q8 6.2 GB ~45 tok/s
RTX 3060 12 GB Gemma 3 12B Q4 11 GB ~35 tok/s
RTX 4060 Ti 16 GB Gemma 3 12B Q6 13 GB ~21 tok/s
RTX 4090 Gemma 3 27B Q4 21 GB ~44 tok/s
RTX A6000 Llama 4 Scout (MoE) Q2 46 GB ~50 tok/s
Mac · M4 Max, 64 GB Gemma 3 27B Q8 33 GB ~14 tok/s

All vision (image input) models, ranked by size

  1. 1. Llama 4 Scout (MoE) 109B · Llama 4 Community

    A 109B MoE with only 17B active and a huge context. Wants ~60 GB+ — a 96–128 GB Mac or multi-GPU box.

  2. 2. Gemma 3 27B 27.4B · Gemma

    The best open multimodal model you can run on one 24 GB card at Q4. Excellent writing, 128K context.

  3. 3. Gemma 3 12B 12.2B · Gemma

    Multimodal, strong writing, 128K context. The 12 GB-card pick when you want vision and long context.

  4. 4. Gemma 3 4B 4.3B · Gemma

    Small and multimodal — it can read images. A good edge pick when you need vision, not just text.

FAQ

What is the best local LLM for vision (image input) right now?

Llama 4 Scout (MoE) is our top pick for vision (image input): A 109B MoE with only 17B active and a huge context. Wants ~60 GB+ — a 96–128 GB Mac or multi-GPU box.

What can I run for vision (image input) on a 24 GB GPU?

On a 24 GB card like the RTX 4090, Gemma 3 27B (Q4) is the strongest vision (image input) model that fits well.

Rankings use parameter count as a capability proxy and our computed fit — see methodology. Pick by benchmark for your exact task.