NVIDIA · 2020-11

Best local LLMs for the A100 80 GB

80 GB VRAM · 2039 GB/s · ~79 GB usable.

Check A100 80 GB price →

Try it with your context & use case

Preset to the A100 80 GB. Change the context length or filter by use case.

28 models run well on A100 80 GB (any context — just fitting the weights) — 79 GB usable.

Sort:
ModelSizeFitBest quantNeedsMemorySpeed
Llama 4 Scout (MoE)MoE109BRuns wellQ5_K_M76 GB
~79 tok/s
Qwen2.5 72B72.7BRuns wellQ8_076 GB
~19 tok/s
Llama 3.3 70B70.6BRuns wellQ8_073 GB
~20 tok/s
DeepSeek-R1 Distill Llama 70B70.6BRuns wellQ8_073 GB
~20 tok/s
Mixtral 8x7B (MoE)MoE46.7BRuns wellQ8_049 GB
~70 tok/s
Qwen3 32B32.8BRuns wellFP1664 GB
~22 tok/s
Qwen2.5 32B32.8BRuns wellFP1664 GB
~22 tok/s
Qwen2.5-Coder 32B32.8BRuns wellFP1664 GB
~22 tok/s
DeepSeek-R1 Distill Qwen 32B32.8BRuns wellFP1664 GB
~22 tok/s
Qwen3 30B-A3B (MoE)MoE30.5BRuns wellFP1660 GB
~145 tok/s
Gemma 3 27B27.4BRuns wellFP1654 GB
~27 tok/s
Gemma 2 27B27.2BRuns wellFP1653 GB
~27 tok/s
Mistral Small 3 24B23.6BRuns wellFP1646 GB
~31 tok/s
Qwen3 14B14.8BRuns wellFP1629 GB
~50 tok/s
Qwen2.5 14B14.8BRuns wellFP1629 GB
~50 tok/s
DeepSeek-R1 Distill Qwen 14B14.8BRuns wellFP1629 GB
~50 tok/s
Phi-4 14B14.7BRuns wellFP1629 GB
~50 tok/s
Mistral Nemo 12B12.2BRuns wellFP1624 GB
~60 tok/s
Gemma 3 12B12.2BRuns wellFP1624 GB
~60 tok/s
Gemma 2 9B9.2BRuns wellFP1619 GB
~79 tok/s
Qwen3 8B8.2BRuns wellFP1617 GB
~90 tok/s
Llama 3.1 8B8.0BRuns wellFP1616 GB
~91 tok/s
Qwen2.5-Coder 7B7.6BRuns wellFP1615 GB
~97 tok/s
Mistral 7B v0.37.3BRuns wellFP1615 GB
~101 tok/s
Gemma 3 4B4.3BRuns wellFP169.1 GB
~171 tok/s
Qwen3 4B4BRuns wellFP168.5 GB
~184 tok/s
Llama 3.2 3B3.2BRuns wellFP167.0 GB
~229 tok/s
Llama 3.2 1B1.2BRuns wellFP163.2 GB
~592 tok/s
DeepSeek-R1 671B-A37B (MoE)MoE671BWon't fit—273 GB
—
Qwen3 235B-A22B (MoE)MoE235BWon't fit—96 GB
—

Estimates, computed in your browser — VRAM, quantisation and speed vary with your runtime and settings. How we estimate →

Top picks for the A100 80 GB

Best for general chat & assistance

Llama 4 Scout (MoE)

109B · Q5 · ~79 tok/s

Best for coding

Qwen2.5 72B

72.7B · Q8 · ~19 tok/s

Best for reasoning & math

Qwen2.5 72B

72.7B · Q8 · ~19 tok/s

Best for writing

Llama 3.3 70B

70.6B · Q8 · ~20 tok/s

Best for vision (image input)

Llama 4 Scout (MoE)

109B · Q5 · ~79 tok/s

Best for low-end & edge hardware

Gemma 3 4B

4.3B · FP16 · ~171 tok/s

Best for tool use & agents

Llama 4 Scout (MoE)

109B · Q5 · ~79 tok/s

Every model on the A100 80 GB

ModelSizeFitBest quantNeedsSpeed
Llama 4 Scout (MoE)MoE 109B Runs well Q5 77 GB ~79 tok/s
Qwen2.5 72B 72.7B Runs well Q8 78 GB ~19 tok/s
Llama 3.3 70B 70.6B Runs well Q8 76 GB ~20 tok/s
DeepSeek-R1 Distill Llama 70B 70.6B Runs well Q8 76 GB ~20 tok/s
Mixtral 8x7B (MoE)MoE 46.7B Runs well Q8 50 GB ~70 tok/s
Qwen3 32B 32.8B Runs well FP16 66 GB ~22 tok/s
Qwen2.5 32B 32.8B Runs well FP16 66 GB ~22 tok/s
Qwen2.5-Coder 32B 32.8B Runs well FP16 66 GB ~22 tok/s
DeepSeek-R1 Distill Qwen 32B 32.8B Runs well FP16 66 GB ~22 tok/s
Qwen3 30B-A3B (MoE)MoE 30.5B Runs well FP16 61 GB ~145 tok/s
Gemma 3 27B 27.4B Runs well FP16 58 GB ~27 tok/s
Gemma 2 27B 27.2B Runs well FP16 56 GB ~27 tok/s
Mistral Small 3 24B 23.6B Runs well FP16 48 GB ~31 tok/s
Qwen3 14B 14.8B Runs well FP16 31 GB ~50 tok/s
Qwen2.5 14B 14.8B Runs well FP16 31 GB ~50 tok/s
DeepSeek-R1 Distill Qwen 14B 14.8B Runs well FP16 31 GB ~50 tok/s
Phi-4 14B 14.7B Runs well FP16 31 GB ~50 tok/s
Mistral Nemo 12B 12.2B Runs well FP16 26 GB ~60 tok/s
Gemma 3 12B 12.2B Runs well FP16 27 GB ~60 tok/s
Gemma 2 9B 9.2B Runs well FP16 21 GB ~79 tok/s
Qwen3 8B 8.2B Runs well FP16 18 GB ~90 tok/s
Llama 3.1 8B 8.0B Runs well FP16 17 GB ~91 tok/s
Qwen2.5-Coder 7B 7.6B Runs well FP16 16 GB ~97 tok/s
Mistral 7B v0.3 7.3B Runs well FP16 16 GB ~101 tok/s
Gemma 3 4B 4.3B Runs well FP16 10 GB ~171 tok/s
Qwen3 4B 4B Runs well FP16 9.6 GB ~184 tok/s
Llama 3.2 3B 3.2B Runs well FP16 7.8 GB ~229 tok/s
Llama 3.2 1B 1.2B Runs well FP16 3.4 GB ~592 tok/s
DeepSeek-R1 671B-A37B (MoE)MoE 671B Won't fit — 303 GB —
Qwen3 235B-A22B (MoE)MoE 235B Won't fit — 98 GB —

Similar hardware

FAQ

What is the best LLM for the A100 80 GB?

For general use, Llama 4 Scout (MoE) is the strongest model that runs well on the A100 80 GB. See the picks-by-use-case below for coding, reasoning and more.

How much can the A100 80 GB run?

The A100 80 GB has 80 GB of VRAM, of which about 79 GB is usable for a model. That runs 28 of our tracked models well.

Estimates at 8K context — see how we compute these.