AMD · 2022-12

Best local LLMs for the Radeon RX 7900 XTX

24 GB VRAM · 960 GB/s · ~23 GB usable. 24 GB at a lower price than NVIDIA — software support is improving but still trails CUDA.

Check Radeon RX 7900 XTX price →

Try it with your context & use case

Preset to the Radeon RX 7900 XTX. Change the context length or filter by use case.

23 models run well and 1 run tight on Radeon RX 7900 XTX (any context — just fitting the weights) — 23 GB usable.

Sort:
ModelSizeFitBest quantNeedsMemorySpeed
Qwen3 32B32.8BRuns wellQ5_K_M23 GB
~30 tok/s
Qwen2.5 32B32.8BRuns wellQ5_K_M23 GB
~30 tok/s
Qwen2.5-Coder 32B32.8BRuns wellQ5_K_M23 GB
~30 tok/s
DeepSeek-R1 Distill Qwen 32B32.8BRuns wellQ5_K_M23 GB
~30 tok/s
Qwen3 30B-A3B (MoE)MoE30.5BRuns wellQ5_K_M22 GB
~192 tok/s
Gemma 3 27B27.4BRuns wellQ6_K23 GB
~31 tok/s
Gemma 2 27B27.2BRuns wellQ6_K22 GB
~31 tok/s
Mistral Small 3 24B23.6BRuns wellQ6_K19 GB
~36 tok/s
Qwen3 14B14.8BRuns wellQ8_016 GB
~44 tok/s
Qwen2.5 14B14.8BRuns wellQ8_016 GB
~44 tok/s
DeepSeek-R1 Distill Qwen 14B14.8BRuns wellQ8_016 GB
~44 tok/s
Phi-4 14B14.7BRuns wellQ8_016 GB
~44 tok/s
Mistral Nemo 12B12.2BRuns wellQ8_013 GB
~53 tok/s
Gemma 3 12B12.2BRuns wellQ8_013 GB
~53 tok/s
Gemma 2 9B9.2BRuns wellFP1619 GB
~37 tok/s
Qwen3 8B8.2BRuns wellFP1617 GB
~42 tok/s
Llama 3.1 8B8.0BRuns wellFP1616 GB
~43 tok/s
Qwen2.5-Coder 7B7.6BRuns wellFP1615 GB
~45 tok/s
Mistral 7B v0.37.3BRuns wellFP1615 GB
~48 tok/s
Gemma 3 4B4.3BRuns wellFP169.1 GB
~80 tok/s
Qwen3 4B4BRuns wellFP168.5 GB
~86 tok/s
Llama 3.2 3B3.2BRuns wellFP167.0 GB
~108 tok/s
Llama 3.2 1B1.2BRuns wellFP163.2 GB
~279 tok/s
Mixtral 8x7B (MoE)MoE46.7BRuns (tight)Q3_K_M23 GB
~71 tok/s
DeepSeek-R1 671B-A37B (MoE)MoE671BWon't fit—273 GB
—
Qwen3 235B-A22B (MoE)MoE235BWon't fit—96 GB
—
Llama 4 Scout (MoE)MoE109BWon't fit—45 GB
—
Qwen2.5 72B72.7BWon't fit—30 GB
—
Llama 3.3 70B70.6BWon't fit—29 GB
—
DeepSeek-R1 Distill Llama 70B70.6BWon't fit—29 GB
—

Estimates, computed in your browser — VRAM, quantisation and speed vary with your runtime and settings. How we estimate →

Top picks for the Radeon RX 7900 XTX

Best for general chat & assistance

Qwen3 32B

32.8B · Q4 · ~35 tok/s

Best for coding

Qwen3 32B

32.8B · Q4 · ~35 tok/s

Best for reasoning & math

Qwen3 32B

32.8B · Q4 · ~35 tok/s

Best for writing

Qwen2.5 32B

32.8B · Q4 · ~35 tok/s

Best for vision (image input)

Gemma 3 27B

27.4B · Q4 · ~42 tok/s

Best for low-end & edge hardware

Gemma 3 4B

4.3B · FP16 · ~80 tok/s

Best for tool use & agents

Qwen3 32B

32.8B · Q4 · ~35 tok/s

Every model on the Radeon RX 7900 XTX

ModelSizeFitBest quantNeedsSpeed
Qwen3 32B 32.8B Runs well Q4 22 GB ~35 tok/s
Qwen2.5 32B 32.8B Runs well Q4 22 GB ~35 tok/s
Qwen2.5-Coder 32B 32.8B Runs well Q4 22 GB ~35 tok/s
DeepSeek-R1 Distill Qwen 32B 32.8B Runs well Q4 22 GB ~35 tok/s
Qwen3 30B-A3B (MoE)MoE 30.5B Runs well Q5 22 GB ~192 tok/s
Gemma 3 27B 27.4B Runs well Q4 21 GB ~42 tok/s
Gemma 2 27B 27.2B Runs well Q5 22 GB ~36 tok/s
Mistral Small 3 24B 23.6B Runs well Q6 21 GB ~36 tok/s
Qwen3 14B 14.8B Runs well Q8 17 GB ~44 tok/s
Qwen2.5 14B 14.8B Runs well Q8 17 GB ~44 tok/s
DeepSeek-R1 Distill Qwen 14B 14.8B Runs well Q8 17 GB ~44 tok/s
Phi-4 14B 14.7B Runs well Q8 17 GB ~44 tok/s
Mistral Nemo 12B 12.2B Runs well Q8 15 GB ~53 tok/s
Gemma 3 12B 12.2B Runs well Q8 16 GB ~53 tok/s
Gemma 2 9B 9.2B Runs well FP16 21 GB ~37 tok/s
Qwen3 8B 8.2B Runs well FP16 18 GB ~42 tok/s
Llama 3.1 8B 8.0B Runs well FP16 17 GB ~43 tok/s
Qwen2.5-Coder 7B 7.6B Runs well FP16 16 GB ~45 tok/s
Mistral 7B v0.3 7.3B Runs well FP16 16 GB ~48 tok/s
Gemma 3 4B 4.3B Runs well FP16 10 GB ~80 tok/s
Qwen3 4B 4B Runs well FP16 9.6 GB ~86 tok/s
Llama 3.2 3B 3.2B Runs well FP16 7.8 GB ~108 tok/s
Llama 3.2 1B 1.2B Runs well FP16 3.4 GB ~279 tok/s
Mixtral 8x7B (MoE)MoE 46.7B Runs (tight) Q2 21 GB ~83 tok/s
DeepSeek-R1 671B-A37B (MoE)MoE 671B Won't fit — 303 GB —
Qwen3 235B-A22B (MoE)MoE 235B Won't fit — 98 GB —
Llama 4 Scout (MoE)MoE 109B Won't fit — 46 GB —
Qwen2.5 72B 72.7B Won't fit — 33 GB —
Llama 3.3 70B 70.6B Won't fit — 32 GB —
DeepSeek-R1 Distill Llama 70B 70.6B Won't fit — 32 GB —

Similar hardware

FAQ

What is the best LLM for the Radeon RX 7900 XTX?

For general use, Qwen3 32B is the strongest model that runs well on the Radeon RX 7900 XTX. See the picks-by-use-case below for coding, reasoning and more.

How much can the Radeon RX 7900 XTX run?

The Radeon RX 7900 XTX has 24 GB of VRAM, of which about 23 GB is usable for a model. That runs 23 of our tracked models well and 1 more at a tight quantisation.

Estimates at 8K context — see how we compute these.