NVIDIA · 2020-09
Best local LLMs for the RTX 3080 10 GB
10 GB VRAM · 760 GB/s · ~9.4 GB usable.
Check RTX 3080 10 GB price →Try it with your context & use case
Preset to the RTX 3080 10 GB. Change the context length or filter by use case.
12 models run well and 3 run tight on RTX 3080 10 GB (any context — just fitting the weights) — 9.4 GB usable.
Sort:
| Model | Size | Fit | Best quant | Needs | Memory | Speed |
|---|---|---|---|---|---|---|
| Phi-4 14B | 14.7B | Runs well | Q4_K_M | 9.3 GB | ~62 tok/s | |
| Mistral Nemo 12B | 12.2B | Runs well | Q5_K_M | 9.1 GB | ~63 tok/s | |
| Gemma 3 12B | 12.2B | Runs well | Q5_K_M | 9.1 GB | ~63 tok/s | |
| Gemma 2 9B | 9.2B | Runs well | Q6_K | 8.1 GB | ~72 tok/s | |
| Qwen3 8B | 8.2B | Runs well | Q8_0 | 9.2 GB | ~63 tok/s | |
| Llama 3.1 8B | 8.0B | Runs well | Q8_0 | 9.0 GB | ~64 tok/s | |
| Qwen2.5-Coder 7B | 7.6B | Runs well | Q8_0 | 8.6 GB | ~68 tok/s | |
| Mistral 7B v0.3 | 7.3B | Runs well | Q8_0 | 8.2 GB | ~71 tok/s | |
| Gemma 3 4B | 4.3B | Runs well | FP16 | 9.1 GB | ~64 tok/s | |
| Qwen3 4B | 4B | Runs well | FP16 | 8.5 GB | ~68 tok/s | |
| Llama 3.2 3B | 3.2B | Runs well | FP16 | 7.0 GB | ~85 tok/s | |
| Llama 3.2 1B | 1.2B | Runs well | FP16 | 3.2 GB | ~221 tok/s | |
| Qwen3 14B | 14.8B | Runs (tight) | Q3_K_M | 7.8 GB | ~76 tok/s | |
| Qwen2.5 14B | 14.8B | Runs (tight) | Q3_K_M | 7.8 GB | ~76 tok/s | |
| DeepSeek-R1 Distill Qwen 14B | 14.8B | Runs (tight) | Q3_K_M | 7.8 GB | ~76 tok/s | |
| DeepSeek-R1 671B-A37B (MoE)MoE | 671B | Won't fit | — | 273 GB | — | |
| Qwen3 235B-A22B (MoE)MoE | 235B | Won't fit | — | 96 GB | — | |
| Llama 4 Scout (MoE)MoE | 109B | Won't fit | — | 45 GB | — | |
| Qwen2.5 72B | 72.7B | Won't fit | — | 30 GB | — | |
| Llama 3.3 70B | 70.6B | Won't fit | — | 29 GB | — | |
| DeepSeek-R1 Distill Llama 70B | 70.6B | Won't fit | — | 29 GB | — | |
| Mixtral 8x7B (MoE)MoE | 46.7B | Won't fit | — | 20 GB | — | |
| Qwen3 32B | 32.8B | Won't fit | — | 14 GB | — | |
| Qwen2.5 32B | 32.8B | Won't fit | — | 14 GB | — | |
| Qwen2.5-Coder 32B | 32.8B | Won't fit | — | 14 GB | — | |
| DeepSeek-R1 Distill Qwen 32B | 32.8B | Won't fit | — | 14 GB | — | |
| Qwen3 30B-A3B (MoE)MoE | 30.5B | Won't fit | — | 13 GB | — | |
| Gemma 3 27B | 27.4B | Won't fit | — | 12 GB | — | |
| Gemma 2 27B | 27.2B | Won't fit | — | 12 GB | — | |
| Mistral Small 3 24B | 23.6B | Won't fit | — | 10 GB | — |
Estimates, computed in your browser — VRAM, quantisation and speed vary with your runtime and settings. How we estimate →
Top picks for the RTX 3080 10 GB
Every model on the RTX 3080 10 GB
| Model | Size | Fit | Best quant | Needs | Speed |
|---|---|---|---|---|---|
| Mistral Nemo 12B | 12.2B | Runs well | Q4 | 9.1 GB | ~74 tok/s |
| Gemma 2 9B | 9.2B | Runs well | Q4 | 8.8 GB | ~98 tok/s |
| Qwen3 8B | 8.2B | Runs well | Q6 | 8.4 GB | ~81 tok/s |
| Llama 3.1 8B | 8.0B | Runs well | Q6 | 8.1 GB | ~83 tok/s |
| Qwen2.5-Coder 7B | 7.6B | Runs well | Q8 | 9.0 GB | ~68 tok/s |
| Mistral 7B v0.3 | 7.3B | Runs well | Q8 | 9.2 GB | ~71 tok/s |
| Gemma 3 4B | 4.3B | Runs well | Q8 | 6.2 GB | ~120 tok/s |
| Qwen3 4B | 4B | Runs well | Q8 | 6.0 GB | ~129 tok/s |
| Llama 3.2 3B | 3.2B | Runs well | FP16 | 7.8 GB | ~85 tok/s |
| Llama 3.2 1B | 1.2B | Runs well | FP16 | 3.4 GB | ~221 tok/s |
| Qwen3 14B | 14.8B | Runs (tight) | Q3 | 9.0 GB | ~76 tok/s |
| Qwen2.5 14B | 14.8B | Runs (tight) | Q3 | 9.3 GB | ~76 tok/s |
| DeepSeek-R1 Distill Qwen 14B | 14.8B | Runs (tight) | Q3 | 9.3 GB | ~76 tok/s |
| Phi-4 14B | 14.7B | Runs (tight) | Q3 | 9.3 GB | ~76 tok/s |
| Gemma 3 12B | 12.2B | Runs (tight) | Q2 | 8.7 GB | ~107 tok/s |
| DeepSeek-R1 671B-A37B (MoE)MoE | 671B | Won't fit | — | 303 GB | — |
| Qwen3 235B-A22B (MoE)MoE | 235B | Won't fit | — | 98 GB | — |
| Llama 4 Scout (MoE)MoE | 109B | Won't fit | — | 46 GB | — |
| Qwen2.5 72B | 72.7B | Won't fit | — | 33 GB | — |
| Llama 3.3 70B | 70.6B | Won't fit | — | 32 GB | — |
| DeepSeek-R1 Distill Llama 70B | 70.6B | Won't fit | — | 32 GB | — |
| Mixtral 8x7B (MoE)MoE | 46.7B | Won't fit | — | 21 GB | — |
| Qwen3 32B | 32.8B | Won't fit | — | 16 GB | — |
| Qwen2.5 32B | 32.8B | Won't fit | — | 16 GB | — |
| Qwen2.5-Coder 32B | 32.8B | Won't fit | — | 16 GB | — |
| DeepSeek-R1 Distill Qwen 32B | 32.8B | Won't fit | — | 16 GB | — |
| Qwen3 30B-A3B (MoE)MoE | 30.5B | Won't fit | — | 14 GB | — |
| Gemma 3 27B | 27.4B | Won't fit | — | 16 GB | — |
| Gemma 2 27B | 27.2B | Won't fit | — | 15 GB | — |
| Mistral Small 3 24B | 23.6B | Won't fit | — | 12 GB | — |
Similar hardware
FAQ
What is the best LLM for the RTX 3080 10 GB?
For general use, Mistral Nemo 12B is the strongest model that runs well on the RTX 3080 10 GB. See the picks-by-use-case below for coding, reasoning and more.
How much can the RTX 3080 10 GB run?
The RTX 3080 10 GB has 10 GB of VRAM, of which about 9.4 GB is usable for a model. That runs 10 of our tracked models well and 5 more at a tight quantisation.
Estimates at 8K context — see how we compute these.