NVIDIA · 2025-01
Best local LLMs for the RTX 5090
32 GB VRAM · 1792 GB/s · ~31 GB usable. 32 GB consumer flagship — the most local-LLM headroom without going workstation.
Check RTX 5090 price →Try it with your context & use case
Preset to the RTX 5090. Change the context length or filter by use case.
24 models run well and 3 run tight on RTX 5090 (any context — just fitting the weights) — 31 GB usable.
| Model | Size | Fit | Best quant | Needs | Memory | Speed |
|---|---|---|---|---|---|---|
| Mixtral 8x7B (MoE)MoE | 46.7B | Runs well | Q4_K_M | 28 GB | ~108 tok/s | |
| Qwen3 32B | 32.8B | Runs well | Q6_K | 27 GB | ~48 tok/s | |
| Qwen2.5 32B | 32.8B | Runs well | Q6_K | 27 GB | ~48 tok/s | |
| Qwen2.5-Coder 32B | 32.8B | Runs well | Q6_K | 27 GB | ~48 tok/s | |
| DeepSeek-R1 Distill Qwen 32B | 32.8B | Runs well | Q6_K | 27 GB | ~48 tok/s | |
| Qwen3 30B-A3B (MoE)MoE | 30.5B | Runs well | Q6_K | 25 GB | ~310 tok/s | |
| Gemma 3 27B | 27.4B | Runs well | Q8_0 | 29 GB | ~44 tok/s | |
| Gemma 2 27B | 27.2B | Runs well | Q8_0 | 29 GB | ~45 tok/s | |
| Mistral Small 3 24B | 23.6B | Runs well | Q8_0 | 25 GB | ~51 tok/s | |
| Qwen3 14B | 14.8B | Runs well | FP16 | 29 GB | ~44 tok/s | |
| Qwen2.5 14B | 14.8B | Runs well | FP16 | 29 GB | ~44 tok/s | |
| DeepSeek-R1 Distill Qwen 14B | 14.8B | Runs well | FP16 | 29 GB | ~44 tok/s | |
| Phi-4 14B | 14.7B | Runs well | FP16 | 29 GB | ~44 tok/s | |
| Mistral Nemo 12B | 12.2B | Runs well | FP16 | 24 GB | ~53 tok/s | |
| Gemma 3 12B | 12.2B | Runs well | FP16 | 24 GB | ~53 tok/s | |
| Gemma 2 9B | 9.2B | Runs well | FP16 | 19 GB | ~70 tok/s | |
| Qwen3 8B | 8.2B | Runs well | FP16 | 17 GB | ~79 tok/s | |
| Llama 3.1 8B | 8.0B | Runs well | FP16 | 16 GB | ~80 tok/s | |
| Qwen2.5-Coder 7B | 7.6B | Runs well | FP16 | 15 GB | ~85 tok/s | |
| Mistral 7B v0.3 | 7.3B | Runs well | FP16 | 15 GB | ~89 tok/s | |
| Gemma 3 4B | 4.3B | Runs well | FP16 | 9.1 GB | ~150 tok/s | |
| Qwen3 4B | 4B | Runs well | FP16 | 8.5 GB | ~161 tok/s | |
| Llama 3.2 3B | 3.2B | Runs well | FP16 | 7.0 GB | ~201 tok/s | |
| Llama 3.2 1B | 1.2B | Runs well | FP16 | 3.2 GB | ~520 tok/s | |
| Qwen2.5 72B | 72.7B | Runs (tight) | Q2_K | 30 GB | ~42 tok/s | |
| Llama 3.3 70B | 70.6B | Runs (tight) | Q2_K | 29 GB | ~44 tok/s | |
| DeepSeek-R1 Distill Llama 70B | 70.6B | Runs (tight) | Q2_K | 29 GB | ~44 tok/s | |
| DeepSeek-R1 671B-A37B (MoE)MoE | 671B | Won't fit | — | 273 GB | — | |
| Qwen3 235B-A22B (MoE)MoE | 235B | Won't fit | — | 96 GB | — | |
| Llama 4 Scout (MoE)MoE | 109B | Won't fit | — | 45 GB | — |
Estimates, computed in your browser — VRAM, quantisation and speed vary with your runtime and settings. How we estimate →
Top picks for the RTX 5090
Every model on the RTX 5090
| Model | Size | Fit | Best quant | Needs | Speed |
|---|---|---|---|---|---|
| Mixtral 8x7B (MoE)MoE | 46.7B | Runs well | Q4 | 29 GB | ~108 tok/s |
| Qwen3 32B | 32.8B | Runs well | Q6 | 29 GB | ~48 tok/s |
| Qwen2.5 32B | 32.8B | Runs well | Q6 | 29 GB | ~48 tok/s |
| Qwen2.5-Coder 32B | 32.8B | Runs well | Q6 | 29 GB | ~48 tok/s |
| DeepSeek-R1 Distill Qwen 32B | 32.8B | Runs well | Q6 | 29 GB | ~48 tok/s |
| Qwen3 30B-A3B (MoE)MoE | 30.5B | Runs well | Q6 | 26 GB | ~310 tok/s |
| Gemma 3 27B | 27.4B | Runs well | Q6 | 26 GB | ~57 tok/s |
| Gemma 2 27B | 27.2B | Runs well | Q6 | 25 GB | ~58 tok/s |
| Mistral Small 3 24B | 23.6B | Runs well | Q8 | 26 GB | ~51 tok/s |
| Qwen3 14B | 14.8B | Runs well | FP16 | 31 GB | ~44 tok/s |
| Qwen2.5 14B | 14.8B | Runs well | FP16 | 31 GB | ~44 tok/s |
| DeepSeek-R1 Distill Qwen 14B | 14.8B | Runs well | FP16 | 31 GB | ~44 tok/s |
| Phi-4 14B | 14.7B | Runs well | FP16 | 31 GB | ~44 tok/s |
| Mistral Nemo 12B | 12.2B | Runs well | FP16 | 26 GB | ~53 tok/s |
| Gemma 3 12B | 12.2B | Runs well | FP16 | 27 GB | ~53 tok/s |
| Gemma 2 9B | 9.2B | Runs well | FP16 | 21 GB | ~70 tok/s |
| Qwen3 8B | 8.2B | Runs well | FP16 | 18 GB | ~79 tok/s |
| Llama 3.1 8B | 8.0B | Runs well | FP16 | 17 GB | ~80 tok/s |
| Qwen2.5-Coder 7B | 7.6B | Runs well | FP16 | 16 GB | ~85 tok/s |
| Mistral 7B v0.3 | 7.3B | Runs well | FP16 | 16 GB | ~89 tok/s |
| Gemma 3 4B | 4.3B | Runs well | FP16 | 10 GB | ~150 tok/s |
| Qwen3 4B | 4B | Runs well | FP16 | 9.6 GB | ~161 tok/s |
| Llama 3.2 3B | 3.2B | Runs well | FP16 | 7.8 GB | ~201 tok/s |
| Llama 3.2 1B | 1.2B | Runs well | FP16 | 3.4 GB | ~520 tok/s |
| DeepSeek-R1 671B-A37B (MoE)MoE | 671B | Won't fit | — | 303 GB | — |
| Qwen3 235B-A22B (MoE)MoE | 235B | Won't fit | — | 98 GB | — |
| Llama 4 Scout (MoE)MoE | 109B | Won't fit | — | 46 GB | — |
| Qwen2.5 72B | 72.7B | Won't fit | — | 33 GB | — |
| Llama 3.3 70B | 70.6B | Won't fit | — | 32 GB | — |
| DeepSeek-R1 Distill Llama 70B | 70.6B | Won't fit | — | 32 GB | — |
Similar hardware
FAQ
What is the best LLM for the RTX 5090?
For general use, Mixtral 8x7B (MoE) is the strongest model that runs well on the RTX 5090. See the picks-by-use-case below for coding, reasoning and more.
How much can the RTX 5090 run?
The RTX 5090 has 32 GB of VRAM, of which about 31 GB is usable for a model. That runs 24 of our tracked models well.
Estimates at 8K context — see how we compute these.