Llama · Dense · Llama 3.2 Community

Can I run Llama 3.2 1B locally?

The smallest genuinely useful Llama. Runs on almost anything — phones, Raspberry Pi, CPU — for classification and simple chat.

Parameters
1.2B
VRAM at Q4
1.7 GB
Max context
131K
Released
2024-09

Memory needed by quantisation

Total includes weights, an 8K-token KV cache and runtime overhead. Lower quants trade quality for size.

QuantBits/weightWeightsTotal neededQuality
FP16 16 2.3 GB 3.4 GB Full precision. Reference quality, 2× the size of 8-bit.
Q8_0 8.5 1.2 GB 2.3 GB Effectively lossless. The safe choice when it fits.
Q6_K 6.56 0.9 GB 2.0 GB Near-lossless; quality loss is hard to measure.
Q5_K_M 5.67 0.8 GB 1.9 GB Very good. A common sweet spot above Q4.
Q4_K_M 4.83 0.7 GB 1.7 GB The default. Best size/quality trade-off for local use.
Q3_K_M 3.91 0.6 GB 1.6 GB Noticeable degradation; useful to squeeze a size up.
Q2_K 3.35 0.5 GB 1.5 GB Aggressive. Quality drops a lot — last resort to fit.

Which hardware runs Llama 3.2 1B?

Best quantisation that fits each device at 8K context, with a rough speed estimate. Try your exact setup →

HardwareMemoryFitBest quantSpeed
RTX 3060 12 GB 12 GB Runs well FP16 ~105 tok/s
RTX 4060 Ti 8 GB 8 GB Runs well FP16 ~84 tok/s
RTX 4060 Ti 16 GB 16 GB Runs well FP16 ~84 tok/s
RTX 3080 10 GB 10 GB Runs well FP16 ~221 tok/s
RTX 5070 12 GB Runs well FP16 ~195 tok/s
RTX 4070 Super 12 GB Runs well FP16 ~146 tok/s
Radeon RX 7900 XT 20 GB Runs well FP16 ~232 tok/s
RTX 5070 Ti 16 GB Runs well FP16 ~260 tok/s
RTX 4070 Ti Super 16 GB Runs well FP16 ~195 tok/s
RTX 3090 24 GB Runs well FP16 ~272 tok/s
Radeon RX 7900 XTX 24 GB Runs well FP16 ~279 tok/s
RTX 5080 16 GB Runs well FP16 ~279 tok/s
RTX 4080 Super 16 GB Runs well FP16 ~214 tok/s
RTX 4090 24 GB Runs well FP16 ~293 tok/s
RTX 5090 32 GB Runs well FP16 ~520 tok/s
RTX A6000 48 GB Runs well FP16 ~223 tok/s
RTX 6000 Ada 48 GB Runs well FP16 ~279 tok/s
A100 80 GB 80 GB Runs well FP16 ~592 tok/s
H100 80 GB 80 GB Runs well FP16 ~973 tok/s
Mac · M1/M2/M3 (base), 8 GB 8 GB Runs well FP16 ~29 tok/s
Mac · M1/M2/M3 (base), 16 GB 16 GB Runs well FP16 ~29 tok/s
Mac · M4 (base), 24 GB 24 GB Runs well FP16 ~35 tok/s
Mac · M4 Pro, 48 GB 48 GB Runs well FP16 ~79 tok/s
Mac · M1/M2/M3 Max, 32 GB 32 GB Runs well FP16 ~116 tok/s
Mac · M4 Max, 64 GB 64 GB Runs well FP16 ~159 tok/s
Mac · M1/M2/M3 Max, 64 GB 64 GB Runs well FP16 ~116 tok/s
Mac · M3/M4 Max, 128 GB 128 GB Runs well FP16 ~159 tok/s
Mac Studio · M1/M2 Ultra, 128 GB 128 GB Runs well FP16 ~232 tok/s
Mac Studio · M3 Ultra, 256 GB 256 GB Runs well FP16 ~238 tok/s
Mac Studio · M3 Ultra, 512 GB 512 GB Runs well FP16 ~238 tok/s
CPU only · 8 GB RAM 8 GB Runs well FP16 ~15 tok/s
CPU only · 16 GB RAM 16 GB Runs well FP16 ~17 tok/s
CPU only · 32 GB RAM 32 GB Runs well FP16 ~20 tok/s
CPU only · 64 GB RAM 64 GB Runs well FP16 ~23 tok/s
CPU only · 128 GB RAM 128 GB Runs well FP16 ~26 tok/s

GPUs that run Llama 3.2 1B well

The most affordable cards in our list that run it at a good quantisation.

Hardware links are affiliate links — they don't change the recommendation.

What Llama 3.2 1B is good for

Related models

FAQ

How much VRAM does Llama 3.2 1B need?

At Q4_K_M, Llama 3.2 1B needs about 1.7 GB including a 8K-token context and overhead (0.7 GB for the weights alone). Higher quantisation needs more; see the table for every level.

What is the cheapest way to run Llama 3.2 1B?

The smallest device that runs it well in our list is the RTX 4060 Ti 8 GB (8 GB). Anything with at least that much memory should handle it at a usable quantisation.

Is Llama 3.2 1B good for low-end & edge hardware?

The smallest genuinely useful Llama. Runs on almost anything — phones, Raspberry Pi, CPU — for classification and simple chat.

Estimates — see how we compute these. Memory figures assume an 8K context; long-context use needs more.