NVIDIA · 2023-03
Best local LLMs for the H100 80 GB
80 GB VRAM · 3350 GB/s · ~79 GB usable.
Check H100 80 GB price →Try it with your context & use case
Preset to the H100 80 GB. Change the context length or filter by use case.
28 models run well on H100 80 GB (any context — just fitting the weights) — 79 GB usable.
Sort:
| Model | Size | Fit | Best quant | Needs | Memory | Speed |
|---|---|---|---|---|---|---|
| Llama 4 Scout (MoE)MoE | 109B | Runs well | Q5_K_M | 76 GB | ~130 tok/s | |
| Qwen2.5 72B | 72.7B | Runs well | Q8_0 | 76 GB | ~31 tok/s | |
| Llama 3.3 70B | 70.6B | Runs well | Q8_0 | 73 GB | ~32 tok/s | |
| DeepSeek-R1 Distill Llama 70B | 70.6B | Runs well | Q8_0 | 73 GB | ~32 tok/s | |
| Mixtral 8x7B (MoE)MoE | 46.7B | Runs well | Q8_0 | 49 GB | ~114 tok/s | |
| Qwen3 32B | 32.8B | Runs well | FP16 | 64 GB | ~37 tok/s | |
| Qwen2.5 32B | 32.8B | Runs well | FP16 | 64 GB | ~37 tok/s | |
| Qwen2.5-Coder 32B | 32.8B | Runs well | FP16 | 64 GB | ~37 tok/s | |
| DeepSeek-R1 Distill Qwen 32B | 32.8B | Runs well | FP16 | 64 GB | ~37 tok/s | |
| Qwen3 30B-A3B (MoE)MoE | 30.5B | Runs well | FP16 | 60 GB | ~238 tok/s | |
| Gemma 3 27B | 27.4B | Runs well | FP16 | 54 GB | ~44 tok/s | |
| Gemma 2 27B | 27.2B | Runs well | FP16 | 53 GB | ~44 tok/s | |
| Mistral Small 3 24B | 23.6B | Runs well | FP16 | 46 GB | ~51 tok/s | |
| Qwen3 14B | 14.8B | Runs well | FP16 | 29 GB | ~81 tok/s | |
| Qwen2.5 14B | 14.8B | Runs well | FP16 | 29 GB | ~81 tok/s | |
| DeepSeek-R1 Distill Qwen 14B | 14.8B | Runs well | FP16 | 29 GB | ~81 tok/s | |
| Phi-4 14B | 14.7B | Runs well | FP16 | 29 GB | ~82 tok/s | |
| Mistral Nemo 12B | 12.2B | Runs well | FP16 | 24 GB | ~99 tok/s | |
| Gemma 3 12B | 12.2B | Runs well | FP16 | 24 GB | ~99 tok/s | |
| Gemma 2 9B | 9.2B | Runs well | FP16 | 19 GB | ~131 tok/s | |
| Qwen3 8B | 8.2B | Runs well | FP16 | 17 GB | ~147 tok/s | |
| Llama 3.1 8B | 8.0B | Runs well | FP16 | 16 GB | ~150 tok/s | |
| Qwen2.5-Coder 7B | 7.6B | Runs well | FP16 | 15 GB | ~159 tok/s | |
| Mistral 7B v0.3 | 7.3B | Runs well | FP16 | 15 GB | ~166 tok/s | |
| Gemma 3 4B | 4.3B | Runs well | FP16 | 9.1 GB | ~280 tok/s | |
| Qwen3 4B | 4B | Runs well | FP16 | 8.5 GB | ~302 tok/s | |
| Llama 3.2 3B | 3.2B | Runs well | FP16 | 7.0 GB | ~376 tok/s | |
| Llama 3.2 1B | 1.2B | Runs well | FP16 | 3.2 GB | ~973 tok/s | |
| DeepSeek-R1 671B-A37B (MoE)MoE | 671B | Won't fit | — | 273 GB | — | |
| Qwen3 235B-A22B (MoE)MoE | 235B | Won't fit | — | 96 GB | — |
Estimates, computed in your browser — VRAM, quantisation and speed vary with your runtime and settings. How we estimate →
Top picks for the H100 80 GB
Every model on the H100 80 GB
| Model | Size | Fit | Best quant | Needs | Speed |
|---|---|---|---|---|---|
| Llama 4 Scout (MoE)MoE | 109B | Runs well | Q5 | 77 GB | ~130 tok/s |
| Qwen2.5 72B | 72.7B | Runs well | Q8 | 78 GB | ~31 tok/s |
| Llama 3.3 70B | 70.6B | Runs well | Q8 | 76 GB | ~32 tok/s |
| DeepSeek-R1 Distill Llama 70B | 70.6B | Runs well | Q8 | 76 GB | ~32 tok/s |
| Mixtral 8x7B (MoE)MoE | 46.7B | Runs well | Q8 | 50 GB | ~114 tok/s |
| Qwen3 32B | 32.8B | Runs well | FP16 | 66 GB | ~37 tok/s |
| Qwen2.5 32B | 32.8B | Runs well | FP16 | 66 GB | ~37 tok/s |
| Qwen2.5-Coder 32B | 32.8B | Runs well | FP16 | 66 GB | ~37 tok/s |
| DeepSeek-R1 Distill Qwen 32B | 32.8B | Runs well | FP16 | 66 GB | ~37 tok/s |
| Qwen3 30B-A3B (MoE)MoE | 30.5B | Runs well | FP16 | 61 GB | ~238 tok/s |
| Gemma 3 27B | 27.4B | Runs well | FP16 | 58 GB | ~44 tok/s |
| Gemma 2 27B | 27.2B | Runs well | FP16 | 56 GB | ~44 tok/s |
| Mistral Small 3 24B | 23.6B | Runs well | FP16 | 48 GB | ~51 tok/s |
| Qwen3 14B | 14.8B | Runs well | FP16 | 31 GB | ~81 tok/s |
| Qwen2.5 14B | 14.8B | Runs well | FP16 | 31 GB | ~81 tok/s |
| DeepSeek-R1 Distill Qwen 14B | 14.8B | Runs well | FP16 | 31 GB | ~81 tok/s |
| Phi-4 14B | 14.7B | Runs well | FP16 | 31 GB | ~82 tok/s |
| Mistral Nemo 12B | 12.2B | Runs well | FP16 | 26 GB | ~99 tok/s |
| Gemma 3 12B | 12.2B | Runs well | FP16 | 27 GB | ~99 tok/s |
| Gemma 2 9B | 9.2B | Runs well | FP16 | 21 GB | ~131 tok/s |
| Qwen3 8B | 8.2B | Runs well | FP16 | 18 GB | ~147 tok/s |
| Llama 3.1 8B | 8.0B | Runs well | FP16 | 17 GB | ~150 tok/s |
| Qwen2.5-Coder 7B | 7.6B | Runs well | FP16 | 16 GB | ~159 tok/s |
| Mistral 7B v0.3 | 7.3B | Runs well | FP16 | 16 GB | ~166 tok/s |
| Gemma 3 4B | 4.3B | Runs well | FP16 | 10 GB | ~280 tok/s |
| Qwen3 4B | 4B | Runs well | FP16 | 9.6 GB | ~302 tok/s |
| Llama 3.2 3B | 3.2B | Runs well | FP16 | 7.8 GB | ~376 tok/s |
| Llama 3.2 1B | 1.2B | Runs well | FP16 | 3.4 GB | ~973 tok/s |
| DeepSeek-R1 671B-A37B (MoE)MoE | 671B | Won't fit | — | 303 GB | — |
| Qwen3 235B-A22B (MoE)MoE | 235B | Won't fit | — | 98 GB | — |
Similar hardware
FAQ
What is the best LLM for the H100 80 GB?
For general use, Llama 4 Scout (MoE) is the strongest model that runs well on the H100 80 GB. See the picks-by-use-case below for coding, reasoning and more.
How much can the H100 80 GB run?
The H100 80 GB has 80 GB of VRAM, of which about 79 GB is usable for a model. That runs 28 of our tracked models well.
Estimates at 8K context — see how we compute these.