NVIDIA · 2020-09

Best local LLMs for the RTX 3080 10 GB

10 GB VRAM · 760 GB/s · ~9.4 GB usable.

Check RTX 3080 10 GB price →

Try it with your context & use case

Preset to the RTX 3080 10 GB. Change the context length or filter by use case.

12 models run well and 3 run tight on RTX 3080 10 GB (any context — just fitting the weights) — 9.4 GB usable.

Sort:
ModelSizeFitBest quantNeedsMemorySpeed
Phi-4 14B14.7BRuns wellQ4_K_M9.3 GB
~62 tok/s
Mistral Nemo 12B12.2BRuns wellQ5_K_M9.1 GB
~63 tok/s
Gemma 3 12B12.2BRuns wellQ5_K_M9.1 GB
~63 tok/s
Gemma 2 9B9.2BRuns wellQ6_K8.1 GB
~72 tok/s
Qwen3 8B8.2BRuns wellQ8_09.2 GB
~63 tok/s
Llama 3.1 8B8.0BRuns wellQ8_09.0 GB
~64 tok/s
Qwen2.5-Coder 7B7.6BRuns wellQ8_08.6 GB
~68 tok/s
Mistral 7B v0.37.3BRuns wellQ8_08.2 GB
~71 tok/s
Gemma 3 4B4.3BRuns wellFP169.1 GB
~64 tok/s
Qwen3 4B4BRuns wellFP168.5 GB
~68 tok/s
Llama 3.2 3B3.2BRuns wellFP167.0 GB
~85 tok/s
Llama 3.2 1B1.2BRuns wellFP163.2 GB
~221 tok/s
Qwen3 14B14.8BRuns (tight)Q3_K_M7.8 GB
~76 tok/s
Qwen2.5 14B14.8BRuns (tight)Q3_K_M7.8 GB
~76 tok/s
DeepSeek-R1 Distill Qwen 14B14.8BRuns (tight)Q3_K_M7.8 GB
~76 tok/s
DeepSeek-R1 671B-A37B (MoE)MoE671BWon't fit—273 GB
—
Qwen3 235B-A22B (MoE)MoE235BWon't fit—96 GB
—
Llama 4 Scout (MoE)MoE109BWon't fit—45 GB
—
Qwen2.5 72B72.7BWon't fit—30 GB
—
Llama 3.3 70B70.6BWon't fit—29 GB
—
DeepSeek-R1 Distill Llama 70B70.6BWon't fit—29 GB
—
Mixtral 8x7B (MoE)MoE46.7BWon't fit—20 GB
—
Qwen3 32B32.8BWon't fit—14 GB
—
Qwen2.5 32B32.8BWon't fit—14 GB
—
Qwen2.5-Coder 32B32.8BWon't fit—14 GB
—
DeepSeek-R1 Distill Qwen 32B32.8BWon't fit—14 GB
—
Qwen3 30B-A3B (MoE)MoE30.5BWon't fit—13 GB
—
Gemma 3 27B27.4BWon't fit—12 GB
—
Gemma 2 27B27.2BWon't fit—12 GB
—
Mistral Small 3 24B23.6BWon't fit—10 GB
—

Estimates, computed in your browser — VRAM, quantisation and speed vary with your runtime and settings. How we estimate →

Top picks for the RTX 3080 10 GB

Best for general chat & assistance

Mistral Nemo 12B

12.2B · Q4 · ~74 tok/s

Best for coding

Qwen2.5-Coder 7B

7.6B · Q8 · ~68 tok/s

Best for reasoning & math

Qwen3 8B

8.2B · Q6 · ~81 tok/s

Best for writing

Mistral Nemo 12B

12.2B · Q4 · ~74 tok/s

Best for vision (image input)

Gemma 3 4B

4.3B · Q8 · ~120 tok/s

Best for low-end & edge hardware

Gemma 3 4B

4.3B · Q8 · ~120 tok/s

Best for tool use & agents

Qwen3 8B

8.2B · Q6 · ~81 tok/s

Every model on the RTX 3080 10 GB

ModelSizeFitBest quantNeedsSpeed
Mistral Nemo 12B 12.2B Runs well Q4 9.1 GB ~74 tok/s
Gemma 2 9B 9.2B Runs well Q4 8.8 GB ~98 tok/s
Qwen3 8B 8.2B Runs well Q6 8.4 GB ~81 tok/s
Llama 3.1 8B 8.0B Runs well Q6 8.1 GB ~83 tok/s
Qwen2.5-Coder 7B 7.6B Runs well Q8 9.0 GB ~68 tok/s
Mistral 7B v0.3 7.3B Runs well Q8 9.2 GB ~71 tok/s
Gemma 3 4B 4.3B Runs well Q8 6.2 GB ~120 tok/s
Qwen3 4B 4B Runs well Q8 6.0 GB ~129 tok/s
Llama 3.2 3B 3.2B Runs well FP16 7.8 GB ~85 tok/s
Llama 3.2 1B 1.2B Runs well FP16 3.4 GB ~221 tok/s
Qwen3 14B 14.8B Runs (tight) Q3 9.0 GB ~76 tok/s
Qwen2.5 14B 14.8B Runs (tight) Q3 9.3 GB ~76 tok/s
DeepSeek-R1 Distill Qwen 14B 14.8B Runs (tight) Q3 9.3 GB ~76 tok/s
Phi-4 14B 14.7B Runs (tight) Q3 9.3 GB ~76 tok/s
Gemma 3 12B 12.2B Runs (tight) Q2 8.7 GB ~107 tok/s
DeepSeek-R1 671B-A37B (MoE)MoE 671B Won't fit — 303 GB —
Qwen3 235B-A22B (MoE)MoE 235B Won't fit — 98 GB —
Llama 4 Scout (MoE)MoE 109B Won't fit — 46 GB —
Qwen2.5 72B 72.7B Won't fit — 33 GB —
Llama 3.3 70B 70.6B Won't fit — 32 GB —
DeepSeek-R1 Distill Llama 70B 70.6B Won't fit — 32 GB —
Mixtral 8x7B (MoE)MoE 46.7B Won't fit — 21 GB —
Qwen3 32B 32.8B Won't fit — 16 GB —
Qwen2.5 32B 32.8B Won't fit — 16 GB —
Qwen2.5-Coder 32B 32.8B Won't fit — 16 GB —
DeepSeek-R1 Distill Qwen 32B 32.8B Won't fit — 16 GB —
Qwen3 30B-A3B (MoE)MoE 30.5B Won't fit — 14 GB —
Gemma 3 27B 27.4B Won't fit — 16 GB —
Gemma 2 27B 27.2B Won't fit — 15 GB —
Mistral Small 3 24B 23.6B Won't fit — 12 GB —

Similar hardware

FAQ

What is the best LLM for the RTX 3080 10 GB?

For general use, Mistral Nemo 12B is the strongest model that runs well on the RTX 3080 10 GB. See the picks-by-use-case below for coding, reasoning and more.

How much can the RTX 3080 10 GB run?

The RTX 3080 10 GB has 10 GB of VRAM, of which about 9.4 GB is usable for a model. That runs 10 of our tracked models well and 5 more at a tight quantisation.

Estimates at 8K context — see how we compute these.