System RAM
Best local LLMs for the CPU only · 128 GB RAM
128 GB system RAM · 90 GB/s · ~125 GB usable.
Try it with your context & use case
Preset to the CPU only · 128 GB RAM. Change the context length or filter by use case.
28 models run well and 1 run tight on CPU only · 128 GB RAM (any context — just fitting the weights) — 125 GB usable (after OS reserve; CPU speeds are low).
| Model | Size | Fit | Best quant | Needs | Memory | Speed |
|---|---|---|---|---|---|---|
| Llama 4 Scout (MoE)MoE | 109B | Runs well | Q8_0 | 113 GB | ~2.3 tok/s | |
| Qwen2.5 72B | 72.7B | Runs well | Q8_0 | 76 GB | <1 tok/s | |
| Llama 3.3 70B | 70.6B | Runs well | Q8_0 | 73 GB | <1 tok/s | |
| DeepSeek-R1 Distill Llama 70B | 70.6B | Runs well | Q8_0 | 73 GB | <1 tok/s | |
| Mixtral 8x7B (MoE)MoE | 46.7B | Runs well | FP16 | 91 GB | ~1.6 tok/s | |
| Qwen3 32B | 32.8B | Runs well | FP16 | 64 GB | <1 tok/s | |
| Qwen2.5 32B | 32.8B | Runs well | FP16 | 64 GB | <1 tok/s | |
| Qwen2.5-Coder 32B | 32.8B | Runs well | FP16 | 64 GB | <1 tok/s | |
| DeepSeek-R1 Distill Qwen 32B | 32.8B | Runs well | FP16 | 64 GB | <1 tok/s | |
| Qwen3 30B-A3B (MoE)MoE | 30.5B | Runs well | FP16 | 60 GB | ~6.4 tok/s | |
| Gemma 3 27B | 27.4B | Runs well | FP16 | 54 GB | ~1.2 tok/s | |
| Gemma 2 27B | 27.2B | Runs well | FP16 | 53 GB | ~1.2 tok/s | |
| Mistral Small 3 24B | 23.6B | Runs well | FP16 | 46 GB | ~1.4 tok/s | |
| Qwen3 14B | 14.8B | Runs well | FP16 | 29 GB | ~2.2 tok/s | |
| Qwen2.5 14B | 14.8B | Runs well | FP16 | 29 GB | ~2.2 tok/s | |
| DeepSeek-R1 Distill Qwen 14B | 14.8B | Runs well | FP16 | 29 GB | ~2.2 tok/s | |
| Phi-4 14B | 14.7B | Runs well | FP16 | 29 GB | ~2.2 tok/s | |
| Mistral Nemo 12B | 12.2B | Runs well | FP16 | 24 GB | ~2.7 tok/s | |
| Gemma 3 12B | 12.2B | Runs well | FP16 | 24 GB | ~2.7 tok/s | |
| Gemma 2 9B | 9.2B | Runs well | FP16 | 19 GB | ~3.5 tok/s | |
| Qwen3 8B | 8.2B | Runs well | FP16 | 17 GB | ~4.0 tok/s | |
| Llama 3.1 8B | 8.0B | Runs well | FP16 | 16 GB | ~4.0 tok/s | |
| Qwen2.5-Coder 7B | 7.6B | Runs well | FP16 | 15 GB | ~4.3 tok/s | |
| Mistral 7B v0.3 | 7.3B | Runs well | FP16 | 15 GB | ~4.5 tok/s | |
| Gemma 3 4B | 4.3B | Runs well | FP16 | 9.1 GB | ~7.5 tok/s | |
| Qwen3 4B | 4B | Runs well | FP16 | 8.5 GB | ~8.1 tok/s | |
| Llama 3.2 3B | 3.2B | Runs well | FP16 | 7.0 GB | ~10 tok/s | |
| Llama 3.2 1B | 1.2B | Runs well | FP16 | 3.2 GB | ~26 tok/s | |
| Qwen3 235B-A22B (MoE)MoE | 235B | Runs (tight) | Q3_K_M | 112 GB | ~3.9 tok/s | |
| DeepSeek-R1 671B-A37B (MoE)MoE | 671B | Won't fit | — | 273 GB | — |
Estimates, computed in your browser — VRAM, quantisation and speed vary with your runtime and settings. How we estimate →
Top picks for the CPU only · 128 GB RAM
Every model on the CPU only · 128 GB RAM
| Model | Size | Fit | Best quant | Needs | Speed |
|---|---|---|---|---|---|
| Llama 4 Scout (MoE)MoE | 109B | Runs well | Q8 | 114 GB | ~2.3 tok/s |
| Qwen2.5 72B | 72.7B | Runs well | Q8 | 78 GB | <1 tok/s |
| Llama 3.3 70B | 70.6B | Runs well | Q8 | 76 GB | <1 tok/s |
| DeepSeek-R1 Distill Llama 70B | 70.6B | Runs well | Q8 | 76 GB | <1 tok/s |
| Mixtral 8x7B (MoE)MoE | 46.7B | Runs well | FP16 | 92 GB | ~1.6 tok/s |
| Qwen3 32B | 32.8B | Runs well | FP16 | 66 GB | <1 tok/s |
| Qwen2.5 32B | 32.8B | Runs well | FP16 | 66 GB | <1 tok/s |
| Qwen2.5-Coder 32B | 32.8B | Runs well | FP16 | 66 GB | <1 tok/s |
| DeepSeek-R1 Distill Qwen 32B | 32.8B | Runs well | FP16 | 66 GB | <1 tok/s |
| Qwen3 30B-A3B (MoE)MoE | 30.5B | Runs well | FP16 | 61 GB | ~6.4 tok/s |
| Gemma 3 27B | 27.4B | Runs well | FP16 | 58 GB | ~1.2 tok/s |
| Gemma 2 27B | 27.2B | Runs well | FP16 | 56 GB | ~1.2 tok/s |
| Mistral Small 3 24B | 23.6B | Runs well | FP16 | 48 GB | ~1.4 tok/s |
| Qwen3 14B | 14.8B | Runs well | FP16 | 31 GB | ~2.2 tok/s |
| Qwen2.5 14B | 14.8B | Runs well | FP16 | 31 GB | ~2.2 tok/s |
| DeepSeek-R1 Distill Qwen 14B | 14.8B | Runs well | FP16 | 31 GB | ~2.2 tok/s |
| Phi-4 14B | 14.7B | Runs well | FP16 | 31 GB | ~2.2 tok/s |
| Mistral Nemo 12B | 12.2B | Runs well | FP16 | 26 GB | ~2.7 tok/s |
| Gemma 3 12B | 12.2B | Runs well | FP16 | 27 GB | ~2.7 tok/s |
| Gemma 2 9B | 9.2B | Runs well | FP16 | 21 GB | ~3.5 tok/s |
| Qwen3 8B | 8.2B | Runs well | FP16 | 18 GB | ~4.0 tok/s |
| Llama 3.1 8B | 8.0B | Runs well | FP16 | 17 GB | ~4.0 tok/s |
| Qwen2.5-Coder 7B | 7.6B | Runs well | FP16 | 16 GB | ~4.3 tok/s |
| Mistral 7B v0.3 | 7.3B | Runs well | FP16 | 16 GB | ~4.5 tok/s |
| Gemma 3 4B | 4.3B | Runs well | FP16 | 10 GB | ~7.5 tok/s |
| Qwen3 4B | 4B | Runs well | FP16 | 9.6 GB | ~8.1 tok/s |
| Llama 3.2 3B | 3.2B | Runs well | FP16 | 7.8 GB | ~10 tok/s |
| Llama 3.2 1B | 1.2B | Runs well | FP16 | 3.4 GB | ~26 tok/s |
| Qwen3 235B-A22B (MoE)MoE | 235B | Runs (tight) | Q3 | 113 GB | ~3.9 tok/s |
| DeepSeek-R1 671B-A37B (MoE)MoE | 671B | Won't fit | — | 303 GB | — |
Similar hardware
FAQ
What is the best LLM for the CPU only · 128 GB RAM?
For general use, Llama 4 Scout (MoE) is the strongest model that runs well on the CPU only · 128 GB RAM. See the picks-by-use-case below for coding, reasoning and more.
How much can the CPU only · 128 GB RAM run?
The CPU only · 128 GB RAM has 128 GB of system RAM, of which about 125 GB is usable for a model. That runs 28 of our tracked models well and 1 more at a tight quantisation.
Estimates at 8K context — see how we compute these.