NVIDIA · 2022-12

Best local LLMs for the RTX 6000 Ada

48 GB VRAM · 960 GB/s · ~47 GB usable.

Check RTX 6000 Ada price →

Try it with your context & use case

Preset to the RTX 6000 Ada. Change the context length or filter by use case.

27 models run well and 1 run tight on RTX 6000 Ada (any context — just fitting the weights) — 47 GB usable.

Sort:
ModelSizeFitBest quantNeedsMemorySpeed
Qwen2.5 72B72.7BRuns wellQ4_K_M43 GB
~16 tok/s
Llama 3.3 70B70.6BRuns wellQ4_K_M42 GB
~16 tok/s
DeepSeek-R1 Distill Llama 70B70.6BRuns wellQ4_K_M42 GB
~16 tok/s
Mixtral 8x7B (MoE)MoE46.7BRuns wellQ6_K38 GB
~42 tok/s
Qwen3 32B32.8BRuns wellQ8_035 GB
~20 tok/s
Qwen2.5 32B32.8BRuns wellQ8_035 GB
~20 tok/s
Qwen2.5-Coder 32B32.8BRuns wellQ8_035 GB
~20 tok/s
DeepSeek-R1 Distill Qwen 32B32.8BRuns wellQ8_035 GB
~20 tok/s
Qwen3 30B-A3B (MoE)MoE30.5BRuns wellQ8_032 GB
~128 tok/s
Gemma 3 27B27.4BRuns wellQ8_029 GB
~24 tok/s
Gemma 2 27B27.2BRuns wellQ8_029 GB
~24 tok/s
Mistral Small 3 24B23.6BRuns wellFP1646 GB
~15 tok/s
Qwen3 14B14.8BRuns wellFP1629 GB
~23 tok/s
Qwen2.5 14B14.8BRuns wellFP1629 GB
~23 tok/s
DeepSeek-R1 Distill Qwen 14B14.8BRuns wellFP1629 GB
~23 tok/s
Phi-4 14B14.7BRuns wellFP1629 GB
~24 tok/s
Mistral Nemo 12B12.2BRuns wellFP1624 GB
~28 tok/s
Gemma 3 12B12.2BRuns wellFP1624 GB
~28 tok/s
Gemma 2 9B9.2BRuns wellFP1619 GB
~37 tok/s
Qwen3 8B8.2BRuns wellFP1617 GB
~42 tok/s
Llama 3.1 8B8.0BRuns wellFP1616 GB
~43 tok/s
Qwen2.5-Coder 7B7.6BRuns wellFP1615 GB
~45 tok/s
Mistral 7B v0.37.3BRuns wellFP1615 GB
~48 tok/s
Gemma 3 4B4.3BRuns wellFP169.1 GB
~80 tok/s
Qwen3 4B4BRuns wellFP168.5 GB
~86 tok/s
Llama 3.2 3B3.2BRuns wellFP167.0 GB
~108 tok/s
Llama 3.2 1B1.2BRuns wellFP163.2 GB
~279 tok/s
Llama 4 Scout (MoE)MoE109BRuns (tight)Q2_K45 GB
~63 tok/s
DeepSeek-R1 671B-A37B (MoE)MoE671BWon't fit—273 GB
—
Qwen3 235B-A22B (MoE)MoE235BWon't fit—96 GB
—

Estimates, computed in your browser — VRAM, quantisation and speed vary with your runtime and settings. How we estimate →

Top picks for the RTX 6000 Ada

Best for general chat & assistance

Qwen2.5 72B

72.7B · Q4 · ~16 tok/s

Best for coding

Qwen2.5 72B

72.7B · Q4 · ~16 tok/s

Best for reasoning & math

Qwen2.5 72B

72.7B · Q4 · ~16 tok/s

Best for writing

Llama 3.3 70B

70.6B · Q4 · ~16 tok/s

Best for vision (image input)

Gemma 3 27B

27.4B · Q8 · ~24 tok/s

Best for low-end & edge hardware

Gemma 3 4B

4.3B · FP16 · ~80 tok/s

Best for tool use & agents

Llama 3.3 70B

70.6B · Q4 · ~16 tok/s

Every model on the RTX 6000 Ada

ModelSizeFitBest quantNeedsSpeed
Qwen2.5 72B 72.7B Runs well Q4 46 GB ~16 tok/s
Llama 3.3 70B 70.6B Runs well Q4 45 GB ~16 tok/s
DeepSeek-R1 Distill Llama 70B 70.6B Runs well Q4 45 GB ~16 tok/s
Mixtral 8x7B (MoE)MoE 46.7B Runs well Q6 39 GB ~42 tok/s
Qwen3 32B 32.8B Runs well Q8 37 GB ~20 tok/s
Qwen2.5 32B 32.8B Runs well Q8 37 GB ~20 tok/s
Qwen2.5-Coder 32B 32.8B Runs well Q8 37 GB ~20 tok/s
DeepSeek-R1 Distill Qwen 32B 32.8B Runs well Q8 37 GB ~20 tok/s
Qwen3 30B-A3B (MoE)MoE 30.5B Runs well Q8 33 GB ~128 tok/s
Gemma 3 27B 27.4B Runs well Q8 33 GB ~24 tok/s
Gemma 2 27B 27.2B Runs well Q8 32 GB ~24 tok/s
Mistral Small 3 24B 23.6B Runs well Q8 26 GB ~28 tok/s
Qwen3 14B 14.8B Runs well FP16 31 GB ~23 tok/s
Qwen2.5 14B 14.8B Runs well FP16 31 GB ~23 tok/s
DeepSeek-R1 Distill Qwen 14B 14.8B Runs well FP16 31 GB ~23 tok/s
Phi-4 14B 14.7B Runs well FP16 31 GB ~24 tok/s
Mistral Nemo 12B 12.2B Runs well FP16 26 GB ~28 tok/s
Gemma 3 12B 12.2B Runs well FP16 27 GB ~28 tok/s
Gemma 2 9B 9.2B Runs well FP16 21 GB ~37 tok/s
Qwen3 8B 8.2B Runs well FP16 18 GB ~42 tok/s
Llama 3.1 8B 8.0B Runs well FP16 17 GB ~43 tok/s
Qwen2.5-Coder 7B 7.6B Runs well FP16 16 GB ~45 tok/s
Mistral 7B v0.3 7.3B Runs well FP16 16 GB ~48 tok/s
Gemma 3 4B 4.3B Runs well FP16 10 GB ~80 tok/s
Qwen3 4B 4B Runs well FP16 9.6 GB ~86 tok/s
Llama 3.2 3B 3.2B Runs well FP16 7.8 GB ~108 tok/s
Llama 3.2 1B 1.2B Runs well FP16 3.4 GB ~279 tok/s
Llama 4 Scout (MoE)MoE 109B Runs (tight) Q2 46 GB ~63 tok/s
DeepSeek-R1 671B-A37B (MoE)MoE 671B Won't fit — 303 GB —
Qwen3 235B-A22B (MoE)MoE 235B Won't fit — 98 GB —

Similar hardware

FAQ

What is the best LLM for the RTX 6000 Ada?

For general use, Qwen2.5 72B is the strongest model that runs well on the RTX 6000 Ada. See the picks-by-use-case below for coding, reasoning and more.

How much can the RTX 6000 Ada run?

The RTX 6000 Ada has 48 GB of VRAM, of which about 47 GB is usable for a model. That runs 27 of our tracked models well and 1 more at a tight quantisation.

Estimates at 8K context — see how we compute these.