DeepSeek · Dense · Llama 3.3 Community
Can I run DeepSeek-R1 Distill Llama 70B locally?
The best open reasoning you can run locally short of the full R1 — if you have ~40 GB of memory.
- Parameters
- 70.6B
- VRAM at Q4
- 45 GB
- Max context
- 131K
- Released
- 2025-01
Memory needed by quantisation
Total includes weights, an 8K-token KV cache and runtime overhead. Lower quants trade quality for size.
| Quant | Bits/weight | Weights | Total needed | Quality |
|---|---|---|---|---|
| FP16 | 16 | 132 GB | 140 GB | Full precision. Reference quality, 2× the size of 8-bit. |
| Q8_0 | 8.5 | 70 GB | 76 GB | Effectively lossless. The safe choice when it fits. |
| Q6_K | 6.56 | 54 GB | 59 GB | Near-lossless; quality loss is hard to measure. |
| Q5_K_M | 5.67 | 47 GB | 52 GB | Very good. A common sweet spot above Q4. |
| Q4_K_M | 4.83 | 40 GB | 45 GB | The default. Best size/quality trade-off for local use. |
| Q3_K_M | 3.91 | 32 GB | 37 GB | Noticeable degradation; useful to squeeze a size up. |
| Q2_K | 3.35 | 28 GB | 32 GB | Aggressive. Quality drops a lot — last resort to fit. |
Which hardware runs DeepSeek-R1 Distill Llama 70B?
Best quantisation that fits each device at 8K context, with a rough speed estimate. Try your exact setup →
| Hardware | Memory | Fit | Best quant | Speed |
|---|---|---|---|---|
| RTX A6000 | 48 GB | Runs well | Q4 | ~13 tok/s |
| RTX 6000 Ada | 48 GB | Runs well | Q4 | ~16 tok/s |
| A100 80 GB | 80 GB | Runs well | Q8 | ~20 tok/s |
| H100 80 GB | 80 GB | Runs well | Q8 | ~32 tok/s |
| Mac · M4 Max, 64 GB | 64 GB | Runs well | Q4 | ~9.2 tok/s |
| Mac · M1/M2/M3 Max, 64 GB | 64 GB | Runs well | Q4 | ~6.8 tok/s |
| Mac · M3/M4 Max, 128 GB | 128 GB | Runs well | Q8 | ~5.2 tok/s |
| Mac Studio · M1/M2 Ultra, 128 GB | 128 GB | Runs well | Q8 | ~7.7 tok/s |
| Mac Studio · M3 Ultra, 256 GB | 256 GB | Runs well | FP16 | ~4.2 tok/s |
| Mac Studio · M3 Ultra, 512 GB | 512 GB | Runs well | FP16 | ~4.2 tok/s |
| CPU only · 64 GB RAM | 64 GB | Runs well | Q6 | <1 tok/s |
| CPU only · 128 GB RAM | 128 GB | Runs well | Q8 | <1 tok/s |
| Mac · M4 Pro, 48 GB | 48 GB | Runs (tight) | Q2 | ~6.6 tok/s |
| RTX 3060 12 GB | 12 GB | Won't fit | — | — |
| RTX 4060 Ti 8 GB | 8 GB | Won't fit | — | — |
| RTX 4060 Ti 16 GB | 16 GB | Won't fit | — | — |
| RTX 3080 10 GB | 10 GB | Won't fit | — | — |
| RTX 5070 | 12 GB | Won't fit | — | — |
| RTX 4070 Super | 12 GB | Won't fit | — | — |
| Radeon RX 7900 XT | 20 GB | Won't fit | — | — |
| RTX 5070 Ti | 16 GB | Won't fit | — | — |
| RTX 4070 Ti Super | 16 GB | Won't fit | — | — |
| RTX 3090 | 24 GB | Won't fit | — | — |
| Radeon RX 7900 XTX | 24 GB | Won't fit | — | — |
| RTX 5080 | 16 GB | Won't fit | — | — |
| RTX 4080 Super | 16 GB | Won't fit | — | — |
| RTX 4090 | 24 GB | Won't fit | — | — |
| RTX 5090 | 32 GB | Won't fit | — | — |
| Mac · M1/M2/M3 (base), 8 GB | 8 GB | Won't fit | — | — |
| Mac · M1/M2/M3 (base), 16 GB | 16 GB | Won't fit | — | — |
| Mac · M4 (base), 24 GB | 24 GB | Won't fit | — | — |
| Mac · M1/M2/M3 Max, 32 GB | 32 GB | Won't fit | — | — |
| CPU only · 8 GB RAM | 8 GB | Won't fit | — | — |
| CPU only · 16 GB RAM | 16 GB | Won't fit | — | — |
| CPU only · 32 GB RAM | 32 GB | Won't fit | — | — |
GPUs that run DeepSeek-R1 Distill Llama 70B well
The most affordable cards in our list that run it at a good quantisation.
RTX A6000
48 GB · ~$4,000
Check price →RTX 6000 Ada
48 GB · ~$6,800
Check price →A100 80 GB
80 GB · ~$15,000
Check price →Hardware links are affiliate links — they don't change the recommendation.
What DeepSeek-R1 Distill Llama 70B is good for
Related models
FAQ
How much VRAM does DeepSeek-R1 Distill Llama 70B need?
At Q4_K_M, DeepSeek-R1 Distill Llama 70B needs about 45 GB including a 8K-token context and overhead (40 GB for the weights alone). Higher quantisation needs more; see the table for every level.
What is the cheapest way to run DeepSeek-R1 Distill Llama 70B?
The smallest device that runs it well in our list is the RTX 6000 Ada (48 GB). Anything with at least that much memory should handle it at a usable quantisation.
Is DeepSeek-R1 Distill Llama 70B good for reasoning & math?
The best open reasoning you can run locally short of the full R1 — if you have ~40 GB of memory.
Estimates — see how we compute these. Memory figures assume an 8K context; long-context use needs more.