Qwen · Mixture-of-experts · Apache 2.0
Can I run Qwen3 30B-A3B (MoE) locally?
Holds 30B in memory but only computes 3B per token — 14B-class quality at near-8B speed. Loves Macs.
- Parameters
- 30.5B (3.3B active)
- VRAM at Q4
- 19 GB
- Max context
- 131K
- Released
- 2025-04
Memory needed by quantisation
Total includes weights, an 8K-token KV cache and runtime overhead. Lower quants trade quality for size.
| Quant | Bits/weight | Weights | Total needed | Quality |
|---|---|---|---|---|
| FP16 | 16 | 57 GB | 61 GB | Full precision. Reference quality, 2× the size of 8-bit. |
| Q8_0 | 8.5 | 30 GB | 33 GB | Effectively lossless. The safe choice when it fits. |
| Q6_K | 6.56 | 23 GB | 26 GB | Near-lossless; quality loss is hard to measure. |
| Q5_K_M | 5.67 | 20 GB | 22 GB | Very good. A common sweet spot above Q4. |
| Q4_K_M | 4.83 | 17 GB | 19 GB | The default. Best size/quality trade-off for local use. |
| Q3_K_M | 3.91 | 14 GB | 16 GB | Noticeable degradation; useful to squeeze a size up. |
| Q2_K | 3.35 | 12 GB | 14 GB | Aggressive. Quality drops a lot — last resort to fit. |
Which hardware runs Qwen3 30B-A3B (MoE)?
Best quantisation that fits each device at 8K context, with a rough speed estimate. Try your exact setup →
| Hardware | Memory | Fit | Best quant | Speed |
|---|---|---|---|---|
| Radeon RX 7900 XT | 20 GB | Runs well | Q4 | ~188 tok/s |
| RTX 3090 | 24 GB | Runs well | Q5 | ~187 tok/s |
| Radeon RX 7900 XTX | 24 GB | Runs well | Q5 | ~192 tok/s |
| RTX 4090 | 24 GB | Runs well | Q5 | ~202 tok/s |
| RTX 5090 | 32 GB | Runs well | Q6 | ~310 tok/s |
| RTX A6000 | 48 GB | Runs well | Q8 | ~103 tok/s |
| RTX 6000 Ada | 48 GB | Runs well | Q8 | ~128 tok/s |
| A100 80 GB | 80 GB | Runs well | FP16 | ~145 tok/s |
| H100 80 GB | 80 GB | Runs well | FP16 | ~238 tok/s |
| Mac · M4 Pro, 48 GB | 48 GB | Runs well | Q8 | ~36 tok/s |
| Mac · M1/M2/M3 Max, 32 GB | 32 GB | Runs well | Q5 | ~80 tok/s |
| Mac · M4 Max, 64 GB | 64 GB | Runs well | Q8 | ~73 tok/s |
| Mac · M1/M2/M3 Max, 64 GB | 64 GB | Runs well | Q8 | ~53 tok/s |
| Mac · M3/M4 Max, 128 GB | 128 GB | Runs well | FP16 | ~39 tok/s |
| Mac Studio · M1/M2 Ultra, 128 GB | 128 GB | Runs well | FP16 | ~57 tok/s |
| Mac Studio · M3 Ultra, 256 GB | 256 GB | Runs well | FP16 | ~58 tok/s |
| Mac Studio · M3 Ultra, 512 GB | 512 GB | Runs well | FP16 | ~58 tok/s |
| CPU only · 32 GB RAM | 32 GB | Runs well | Q6 | ~12 tok/s |
| CPU only · 64 GB RAM | 64 GB | Runs well | FP16 | ~5.7 tok/s |
| CPU only · 128 GB RAM | 128 GB | Runs well | FP16 | ~6.4 tok/s |
| RTX 4060 Ti 16 GB | 16 GB | Runs (tight) | Q2 | ~98 tok/s |
| RTX 5070 Ti | 16 GB | Runs (tight) | Q2 | ~303 tok/s |
| RTX 4070 Ti Super | 16 GB | Runs (tight) | Q2 | ~228 tok/s |
| RTX 5080 | 16 GB | Runs (tight) | Q2 | ~325 tok/s |
| RTX 4080 Super | 16 GB | Runs (tight) | Q2 | ~249 tok/s |
| Mac · M4 (base), 24 GB | 24 GB | Runs (tight) | Q3 | ~35 tok/s |
| RTX 3060 12 GB | 12 GB | Won't fit | — | — |
| RTX 4060 Ti 8 GB | 8 GB | Won't fit | — | — |
| RTX 3080 10 GB | 10 GB | Won't fit | — | — |
| RTX 5070 | 12 GB | Won't fit | — | — |
| RTX 4070 Super | 12 GB | Won't fit | — | — |
| Mac · M1/M2/M3 (base), 8 GB | 8 GB | Won't fit | — | — |
| Mac · M1/M2/M3 (base), 16 GB | 16 GB | Won't fit | — | — |
| CPU only · 8 GB RAM | 8 GB | Won't fit | — | — |
| CPU only · 16 GB RAM | 16 GB | Won't fit | — | — |
GPUs that run Qwen3 30B-A3B (MoE) well
The most affordable cards in our list that run it at a good quantisation.
Radeon RX 7900 XT
20 GB · ~$700
Check price →RTX 3090
24 GB · ~$800
Check price →Radeon RX 7900 XTX
24 GB · ~$900
Check price →Hardware links are affiliate links — they don't change the recommendation.
What Qwen3 30B-A3B (MoE) is good for
Related models
FAQ
How much VRAM does Qwen3 30B-A3B (MoE) need?
At Q4_K_M, Qwen3 30B-A3B (MoE) needs about 19 GB including a 8K-token context and overhead (17 GB for the weights alone). Higher quantisation needs more; see the table for every level.
What is the cheapest way to run Qwen3 30B-A3B (MoE)?
The smallest device that runs it well in our list is the Radeon RX 7900 XT (20 GB). Anything with at least that much memory should handle it at a usable quantisation.
Is Qwen3 30B-A3B (MoE) good for general chat & assistance?
Holds 30B in memory but only computes 3B per token — 14B-class quality at near-8B speed. Loves Macs.
Estimates — see how we compute these. Memory figures assume an 8K context; long-context use needs more.