DeepSeek · Dense · MIT

Can I run DeepSeek-R1 Distill Qwen 14B locally?

R1’s reasoning distilled into a 14B that fits a 12 GB card. Pick this for math and step-by-step problems.

Parameters
14.8B
VRAM at Q4
11 GB
Max context
131K
Released
2025-01

Memory needed by quantisation

Total includes weights, an 8K-token KV cache and runtime overhead. Lower quants trade quality for size.

QuantBits/weightWeightsTotal neededQuality
FP16 16 28 GB 31 GB Full precision. Reference quality, 2× the size of 8-bit.
Q8_0 8.5 15 GB 17 GB Effectively lossless. The safe choice when it fits.
Q6_K 6.56 11 GB 14 GB Near-lossless; quality loss is hard to measure.
Q5_K_M 5.67 9.8 GB 12 GB Very good. A common sweet spot above Q4.
Q4_K_M 4.83 8.3 GB 11 GB The default. Best size/quality trade-off for local use.
Q3_K_M 3.91 6.7 GB 9.3 GB Noticeable degradation; useful to squeeze a size up.
Q2_K 3.35 5.8 GB 8.3 GB Aggressive. Quality drops a lot — last resort to fit.

Which hardware runs DeepSeek-R1 Distill Qwen 14B?

Best quantisation that fits each device at 8K context, with a rough speed estimate. Try your exact setup →

HardwareMemoryFitBest quantSpeed
RTX 3060 12 GB 12 GB Runs well Q4 ~29 tok/s
RTX 4060 Ti 16 GB 16 GB Runs well Q6 ~17 tok/s
RTX 5070 12 GB Runs well Q4 ~54 tok/s
RTX 4070 Super 12 GB Runs well Q4 ~41 tok/s
Radeon RX 7900 XT 20 GB Runs well Q8 ~37 tok/s
RTX 5070 Ti 16 GB Runs well Q6 ~53 tok/s
RTX 4070 Ti Super 16 GB Runs well Q6 ~40 tok/s
RTX 3090 24 GB Runs well Q8 ~43 tok/s
Radeon RX 7900 XTX 24 GB Runs well Q8 ~44 tok/s
RTX 5080 16 GB Runs well Q6 ~57 tok/s
RTX 4080 Super 16 GB Runs well Q6 ~44 tok/s
RTX 4090 24 GB Runs well Q8 ~46 tok/s
RTX 5090 32 GB Runs well FP16 ~44 tok/s
RTX A6000 48 GB Runs well FP16 ~19 tok/s
RTX 6000 Ada 48 GB Runs well FP16 ~23 tok/s
A100 80 GB 80 GB Runs well FP16 ~50 tok/s
H100 80 GB 80 GB Runs well FP16 ~81 tok/s
Mac · M1/M2/M3 (base), 16 GB 16 GB Runs well Q4 ~8.1 tok/s
Mac · M4 (base), 24 GB 24 GB Runs well Q6 ~7.1 tok/s
Mac · M4 Pro, 48 GB 48 GB Runs well FP16 ~6.6 tok/s
Mac · M1/M2/M3 Max, 32 GB 32 GB Runs well Q8 ~18 tok/s
Mac · M4 Max, 64 GB 64 GB Runs well FP16 ~13 tok/s
Mac · M1/M2/M3 Max, 64 GB 64 GB Runs well FP16 ~9.7 tok/s
Mac · M3/M4 Max, 128 GB 128 GB Runs well FP16 ~13 tok/s
Mac Studio · M1/M2 Ultra, 128 GB 128 GB Runs well FP16 ~19 tok/s
Mac Studio · M3 Ultra, 256 GB 256 GB Runs well FP16 ~20 tok/s
Mac Studio · M3 Ultra, 512 GB 512 GB Runs well FP16 ~20 tok/s
CPU only · 16 GB RAM 16 GB Runs well Q5 ~4.1 tok/s
CPU only · 32 GB RAM 32 GB Runs well Q8 ~3.2 tok/s
CPU only · 64 GB RAM 64 GB Runs well FP16 ~1.9 tok/s
CPU only · 128 GB RAM 128 GB Runs well FP16 ~2.2 tok/s
RTX 3080 10 GB 10 GB Runs (tight) Q3 ~76 tok/s
RTX 4060 Ti 8 GB 8 GB Won't fit — —
Mac · M1/M2/M3 (base), 8 GB 8 GB Won't fit — —
CPU only · 8 GB RAM 8 GB Won't fit — —

GPUs that run DeepSeek-R1 Distill Qwen 14B well

The most affordable cards in our list that run it at a good quantisation.

Hardware links are affiliate links — they don't change the recommendation.

What DeepSeek-R1 Distill Qwen 14B is good for

Related models

FAQ

How much VRAM does DeepSeek-R1 Distill Qwen 14B need?

At Q4_K_M, DeepSeek-R1 Distill Qwen 14B needs about 11 GB including a 8K-token context and overhead (8.3 GB for the weights alone). Higher quantisation needs more; see the table for every level.

What is the cheapest way to run DeepSeek-R1 Distill Qwen 14B?

The smallest device that runs it well in our list is the RTX 5070 (12 GB). Anything with at least that much memory should handle it at a usable quantisation.

Is DeepSeek-R1 Distill Qwen 14B good for reasoning & math?

R1’s reasoning distilled into a 14B that fits a 12 GB card. Pick this for math and step-by-step problems.

Estimates — see how we compute these. Memory figures assume an 8K context; long-context use needs more.