Qwen · Dense · Apache 2.0
Can I run Qwen2.5-Coder 7B locally?
The best coding model that fits 8 GB. Strong fill-in-the-middle for editor autocomplete.
- Parameters
- 7.6B
- VRAM at Q4
- 5.6 GB
- Max context
- 131K
- Released
- 2024-11
Memory needed by quantisation
Total includes weights, an 8K-token KV cache and runtime overhead. Lower quants trade quality for size.
| Quant | Bits/weight | Weights | Total needed | Quality |
|---|---|---|---|---|
| FP16 | 16 | 14 GB | 16 GB | Full precision. Reference quality, 2× the size of 8-bit. |
| Q8_0 | 8.5 | 7.5 GB | 9.0 GB | Effectively lossless. The safe choice when it fits. |
| Q6_K | 6.56 | 5.8 GB | 7.2 GB | Near-lossless; quality loss is hard to measure. |
| Q5_K_M | 5.67 | 5.0 GB | 6.4 GB | Very good. A common sweet spot above Q4. |
| Q4_K_M | 4.83 | 4.3 GB | 5.6 GB | The default. Best size/quality trade-off for local use. |
| Q3_K_M | 3.91 | 3.5 GB | 4.8 GB | Noticeable degradation; useful to squeeze a size up. |
| Q2_K | 3.35 | 3.0 GB | 4.3 GB | Aggressive. Quality drops a lot — last resort to fit. |
Which hardware runs Qwen2.5-Coder 7B?
Best quantisation that fits each device at 8K context, with a rough speed estimate. Try your exact setup →
| Hardware | Memory | Fit | Best quant | Speed |
|---|---|---|---|---|
| RTX 3060 12 GB | 12 GB | Runs well | Q8 | ~32 tok/s |
| RTX 4060 Ti 8 GB | 8 GB | Runs well | Q6 | ~33 tok/s |
| RTX 4060 Ti 16 GB | 16 GB | Runs well | Q8 | ~26 tok/s |
| RTX 3080 10 GB | 10 GB | Runs well | Q8 | ~68 tok/s |
| RTX 5070 | 12 GB | Runs well | Q8 | ~60 tok/s |
| RTX 4070 Super | 12 GB | Runs well | Q8 | ~45 tok/s |
| Radeon RX 7900 XT | 20 GB | Runs well | FP16 | ~38 tok/s |
| RTX 5070 Ti | 16 GB | Runs well | Q8 | ~80 tok/s |
| RTX 4070 Ti Super | 16 GB | Runs well | Q8 | ~60 tok/s |
| RTX 3090 | 24 GB | Runs well | FP16 | ~44 tok/s |
| Radeon RX 7900 XTX | 24 GB | Runs well | FP16 | ~45 tok/s |
| RTX 5080 | 16 GB | Runs well | Q8 | ~86 tok/s |
| RTX 4080 Super | 16 GB | Runs well | Q8 | ~66 tok/s |
| RTX 4090 | 24 GB | Runs well | FP16 | ~48 tok/s |
| RTX 5090 | 32 GB | Runs well | FP16 | ~85 tok/s |
| RTX A6000 | 48 GB | Runs well | FP16 | ~36 tok/s |
| RTX 6000 Ada | 48 GB | Runs well | FP16 | ~45 tok/s |
| A100 80 GB | 80 GB | Runs well | FP16 | ~97 tok/s |
| H100 80 GB | 80 GB | Runs well | FP16 | ~159 tok/s |
| Mac · M1/M2/M3 (base), 8 GB | 8 GB | Runs well | Q4 | ~16 tok/s |
| Mac · M1/M2/M3 (base), 16 GB | 16 GB | Runs well | Q8 | ~8.9 tok/s |
| Mac · M4 (base), 24 GB | 24 GB | Runs well | FP16 | ~5.7 tok/s |
| Mac · M4 Pro, 48 GB | 48 GB | Runs well | FP16 | ~13 tok/s |
| Mac · M1/M2/M3 Max, 32 GB | 32 GB | Runs well | FP16 | ~19 tok/s |
| Mac · M4 Max, 64 GB | 64 GB | Runs well | FP16 | ~26 tok/s |
| Mac · M1/M2/M3 Max, 64 GB | 64 GB | Runs well | FP16 | ~19 tok/s |
| Mac · M3/M4 Max, 128 GB | 128 GB | Runs well | FP16 | ~26 tok/s |
| Mac Studio · M1/M2 Ultra, 128 GB | 128 GB | Runs well | FP16 | ~38 tok/s |
| Mac Studio · M3 Ultra, 256 GB | 256 GB | Runs well | FP16 | ~39 tok/s |
| Mac Studio · M3 Ultra, 512 GB | 512 GB | Runs well | FP16 | ~39 tok/s |
| CPU only · 16 GB RAM | 16 GB | Runs well | Q8 | ~5.3 tok/s |
| CPU only · 32 GB RAM | 32 GB | Runs well | FP16 | ~3.3 tok/s |
| CPU only · 64 GB RAM | 64 GB | Runs well | FP16 | ~3.8 tok/s |
| CPU only · 128 GB RAM | 128 GB | Runs well | FP16 | ~4.3 tok/s |
| CPU only · 8 GB RAM | 8 GB | Runs (tight) | Q3 | ~9.7 tok/s |
GPUs that run Qwen2.5-Coder 7B well
The most affordable cards in our list that run it at a good quantisation.
RTX 3060 12 GB
12 GB · ~$279
Check price →RTX 4060 Ti 8 GB
8 GB · ~$379
Check price →RTX 4060 Ti 16 GB
16 GB · ~$449
Check price →Hardware links are affiliate links — they don't change the recommendation.
What Qwen2.5-Coder 7B is good for
Related models
FAQ
How much VRAM does Qwen2.5-Coder 7B need?
At Q4_K_M, Qwen2.5-Coder 7B needs about 5.6 GB including a 8K-token context and overhead (4.3 GB for the weights alone). Higher quantisation needs more; see the table for every level.
What is the cheapest way to run Qwen2.5-Coder 7B?
The smallest device that runs it well in our list is the RTX 4060 Ti 8 GB (8 GB). Anything with at least that much memory should handle it at a usable quantisation.
Is Qwen2.5-Coder 7B good for coding?
The best coding model that fits 8 GB. Strong fill-in-the-middle for editor autocomplete.
Estimates — see how we compute these. Memory figures assume an 8K context; long-context use needs more.