AMD · 2022-12
Best local LLMs for the Radeon RX 7900 XT
20 GB VRAM · 800 GB/s · ~19 GB usable.
Check Radeon RX 7900 XT price →Try it with your context & use case
Preset to the Radeon RX 7900 XT. Change the context length or filter by use case.
19 models run well and 4 run tight on Radeon RX 7900 XT (any context — just fitting the weights) — 19 GB usable.
Sort:
| Model | Size | Fit | Best quant | Needs | Memory | Speed |
|---|---|---|---|---|---|---|
| Qwen3 30B-A3B (MoE)MoE | 30.5B | Runs well | Q4_K_M | 19 GB | ~188 tok/s | |
| Gemma 3 27B | 27.4B | Runs well | Q4_K_M | 17 GB | ~35 tok/s | |
| Gemma 2 27B | 27.2B | Runs well | Q4_K_M | 17 GB | ~35 tok/s | |
| Mistral Small 3 24B | 23.6B | Runs well | Q5_K_M | 17 GB | ~34 tok/s | |
| Qwen3 14B | 14.8B | Runs well | Q8_0 | 16 GB | ~37 tok/s | |
| Qwen2.5 14B | 14.8B | Runs well | Q8_0 | 16 GB | ~37 tok/s | |
| DeepSeek-R1 Distill Qwen 14B | 14.8B | Runs well | Q8_0 | 16 GB | ~37 tok/s | |
| Phi-4 14B | 14.7B | Runs well | Q8_0 | 16 GB | ~37 tok/s | |
| Mistral Nemo 12B | 12.2B | Runs well | Q8_0 | 13 GB | ~44 tok/s | |
| Gemma 3 12B | 12.2B | Runs well | Q8_0 | 13 GB | ~44 tok/s | |
| Gemma 2 9B | 9.2B | Runs well | FP16 | 19 GB | ~31 tok/s | |
| Qwen3 8B | 8.2B | Runs well | FP16 | 17 GB | ~35 tok/s | |
| Llama 3.1 8B | 8.0B | Runs well | FP16 | 16 GB | ~36 tok/s | |
| Qwen2.5-Coder 7B | 7.6B | Runs well | FP16 | 15 GB | ~38 tok/s | |
| Mistral 7B v0.3 | 7.3B | Runs well | FP16 | 15 GB | ~40 tok/s | |
| Gemma 3 4B | 4.3B | Runs well | FP16 | 9.1 GB | ~67 tok/s | |
| Qwen3 4B | 4B | Runs well | FP16 | 8.5 GB | ~72 tok/s | |
| Llama 3.2 3B | 3.2B | Runs well | FP16 | 7.0 GB | ~90 tok/s | |
| Llama 3.2 1B | 1.2B | Runs well | FP16 | 3.2 GB | ~232 tok/s | |
| Qwen3 32B | 32.8B | Runs (tight) | Q3_K_M | 16 GB | ~36 tok/s | |
| Qwen2.5 32B | 32.8B | Runs (tight) | Q3_K_M | 16 GB | ~36 tok/s | |
| Qwen2.5-Coder 32B | 32.8B | Runs (tight) | Q3_K_M | 16 GB | ~36 tok/s | |
| DeepSeek-R1 Distill Qwen 32B | 32.8B | Runs (tight) | Q3_K_M | 16 GB | ~36 tok/s | |
| DeepSeek-R1 671B-A37B (MoE)MoE | 671B | Won't fit | — | 273 GB | — | |
| Qwen3 235B-A22B (MoE)MoE | 235B | Won't fit | — | 96 GB | — | |
| Llama 4 Scout (MoE)MoE | 109B | Won't fit | — | 45 GB | — | |
| Qwen2.5 72B | 72.7B | Won't fit | — | 30 GB | — | |
| Llama 3.3 70B | 70.6B | Won't fit | — | 29 GB | — | |
| DeepSeek-R1 Distill Llama 70B | 70.6B | Won't fit | — | 29 GB | — | |
| Mixtral 8x7B (MoE)MoE | 46.7B | Won't fit | — | 20 GB | — |
Estimates, computed in your browser — VRAM, quantisation and speed vary with your runtime and settings. How we estimate →
Top picks for the Radeon RX 7900 XT
Every model on the Radeon RX 7900 XT
| Model | Size | Fit | Best quant | Needs | Speed |
|---|---|---|---|---|---|
| Qwen3 30B-A3B (MoE)MoE | 30.5B | Runs well | Q4 | 19 GB | ~188 tok/s |
| Mistral Small 3 24B | 23.6B | Runs well | Q5 | 18 GB | ~34 tok/s |
| Qwen3 14B | 14.8B | Runs well | Q8 | 17 GB | ~37 tok/s |
| Qwen2.5 14B | 14.8B | Runs well | Q8 | 17 GB | ~37 tok/s |
| DeepSeek-R1 Distill Qwen 14B | 14.8B | Runs well | Q8 | 17 GB | ~37 tok/s |
| Phi-4 14B | 14.7B | Runs well | Q8 | 17 GB | ~37 tok/s |
| Mistral Nemo 12B | 12.2B | Runs well | Q8 | 15 GB | ~44 tok/s |
| Gemma 3 12B | 12.2B | Runs well | Q8 | 16 GB | ~44 tok/s |
| Gemma 2 9B | 9.2B | Runs well | Q8 | 13 GB | ~59 tok/s |
| Qwen3 8B | 8.2B | Runs well | FP16 | 18 GB | ~35 tok/s |
| Llama 3.1 8B | 8.0B | Runs well | FP16 | 17 GB | ~36 tok/s |
| Qwen2.5-Coder 7B | 7.6B | Runs well | FP16 | 16 GB | ~38 tok/s |
| Mistral 7B v0.3 | 7.3B | Runs well | FP16 | 16 GB | ~40 tok/s |
| Gemma 3 4B | 4.3B | Runs well | FP16 | 10 GB | ~67 tok/s |
| Qwen3 4B | 4B | Runs well | FP16 | 9.6 GB | ~72 tok/s |
| Llama 3.2 3B | 3.2B | Runs well | FP16 | 7.8 GB | ~90 tok/s |
| Llama 3.2 1B | 1.2B | Runs well | FP16 | 3.4 GB | ~232 tok/s |
| Qwen3 32B | 32.8B | Runs (tight) | Q3 | 18 GB | ~36 tok/s |
| Qwen2.5 32B | 32.8B | Runs (tight) | Q3 | 18 GB | ~36 tok/s |
| Qwen2.5-Coder 32B | 32.8B | Runs (tight) | Q3 | 18 GB | ~36 tok/s |
| DeepSeek-R1 Distill Qwen 32B | 32.8B | Runs (tight) | Q3 | 18 GB | ~36 tok/s |
| Gemma 3 27B | 27.4B | Runs (tight) | Q3 | 18 GB | ~43 tok/s |
| Gemma 2 27B | 27.2B | Runs (tight) | Q3 | 17 GB | ~43 tok/s |
| DeepSeek-R1 671B-A37B (MoE)MoE | 671B | Won't fit | — | 303 GB | — |
| Qwen3 235B-A22B (MoE)MoE | 235B | Won't fit | — | 98 GB | — |
| Llama 4 Scout (MoE)MoE | 109B | Won't fit | — | 46 GB | — |
| Qwen2.5 72B | 72.7B | Won't fit | — | 33 GB | — |
| Llama 3.3 70B | 70.6B | Won't fit | — | 32 GB | — |
| DeepSeek-R1 Distill Llama 70B | 70.6B | Won't fit | — | 32 GB | — |
| Mixtral 8x7B (MoE)MoE | 46.7B | Won't fit | — | 21 GB | — |
Similar hardware
FAQ
What is the best LLM for the Radeon RX 7900 XT?
For general use, Qwen3 30B-A3B (MoE) is the strongest model that runs well on the Radeon RX 7900 XT. See the picks-by-use-case below for coding, reasoning and more.
How much can the Radeon RX 7900 XT run?
The Radeon RX 7900 XT has 20 GB of VRAM, of which about 19 GB is usable for a model. That runs 17 of our tracked models well and 6 more at a tight quantisation.
Estimates at 8K context — see how we compute these.