Apple

Best local LLMs for the Mac · M1/M2/M3 (base), 8 GB

8 GB unified memory · 100 GB/s · ~5.8 GB usable. Only ~5–6 GB is usable for a model. Small models only.

Try it with your context & use case

Preset to the Mac · M1/M2/M3 (base), 8 GB. Change the context length or filter by use case.

8 models run well and 3 run tight on Mac · M1/M2/M3 (base), 8 GB (any context — just fitting the weights) — 5.8 GB usable (≈72% of unified memory).

Sort:
ModelSizeFitBest quantNeedsMemorySpeed
Qwen3 8B8.2BRuns wellQ4_K_M5.5 GB
~15 tok/s
Llama 3.1 8B8.0BRuns wellQ4_K_M5.4 GB
~15 tok/s
Qwen2.5-Coder 7B7.6BRuns wellQ4_K_M5.2 GB
~16 tok/s
Mistral 7B v0.37.3BRuns wellQ5_K_M5.7 GB
~14 tok/s
Gemma 3 4B4.3BRuns wellQ8_05.2 GB
~16 tok/s
Qwen3 4B4BRuns wellQ8_04.9 GB
~17 tok/s
Llama 3.2 3B3.2BRuns wellQ8_04.1 GB
~21 tok/s
Llama 3.2 1B1.2BRuns wellFP163.2 GB
~29 tok/s
Mistral Nemo 12B12.2BRuns (tight)Q2_K5.7 GB
~14 tok/s
Gemma 3 12B12.2BRuns (tight)Q2_K5.7 GB
~14 tok/s
Gemma 2 9B9.2BRuns (tight)Q3_K_M5.1 GB
~16 tok/s
DeepSeek-R1 671B-A37B (MoE)MoE671BWon't fit—273 GB
—
Qwen3 235B-A22B (MoE)MoE235BWon't fit—96 GB
—
Llama 4 Scout (MoE)MoE109BWon't fit—45 GB
—
Qwen2.5 72B72.7BWon't fit—30 GB
—
Llama 3.3 70B70.6BWon't fit—29 GB
—
DeepSeek-R1 Distill Llama 70B70.6BWon't fit—29 GB
—
Mixtral 8x7B (MoE)MoE46.7BWon't fit—20 GB
—
Qwen3 32B32.8BWon't fit—14 GB
—
Qwen2.5 32B32.8BWon't fit—14 GB
—
Qwen2.5-Coder 32B32.8BWon't fit—14 GB
—
DeepSeek-R1 Distill Qwen 32B32.8BWon't fit—14 GB
—
Qwen3 30B-A3B (MoE)MoE30.5BWon't fit—13 GB
—
Gemma 3 27B27.4BWon't fit—12 GB
—
Gemma 2 27B27.2BWon't fit—12 GB
—
Mistral Small 3 24B23.6BWon't fit—10 GB
—
Qwen3 14B14.8BWon't fit—6.8 GB
—
Qwen2.5 14B14.8BWon't fit—6.8 GB
—
DeepSeek-R1 Distill Qwen 14B14.8BWon't fit—6.8 GB
—
Phi-4 14B14.7BWon't fit—6.7 GB
—

Estimates, computed in your browser — VRAM, quantisation and speed vary with your runtime and settings. How we estimate →

Top picks for the Mac · M1/M2/M3 (base), 8 GB

Best for general chat & assistance

Gemma 3 4B

4.3B · Q6 · ~20 tok/s

Best for coding

Qwen2.5-Coder 7B

7.6B · Q4 · ~16 tok/s

Best for reasoning & math

Qwen3 4B

4B · Q6 · ~22 tok/s

Best for writing

Llama 3.1 8B

8.0B · Q3 · ~18 tok/s

Best for vision (image input)

Gemma 3 4B

4.3B · Q6 · ~20 tok/s

Best for low-end & edge hardware

Gemma 3 4B

4.3B · Q6 · ~20 tok/s

Best for tool use & agents

Qwen2.5-Coder 7B

7.6B · Q4 · ~16 tok/s

Every model on the Mac · M1/M2/M3 (base), 8 GB

ModelSizeFitBest quantNeedsSpeed
Qwen2.5-Coder 7B 7.6B Runs well Q4 5.6 GB ~16 tok/s
Gemma 3 4B 4.3B Runs well Q6 5.2 GB ~20 tok/s
Qwen3 4B 4B Runs well Q6 5.1 GB ~22 tok/s
Llama 3.2 3B 3.2B Runs well Q8 4.9 GB ~21 tok/s
Llama 3.2 1B 1.2B Runs well FP16 3.4 GB ~29 tok/s
Qwen3 8B 8.2B Runs (tight) Q3 5.8 GB ~18 tok/s
Llama 3.1 8B 8.0B Runs (tight) Q3 5.6 GB ~18 tok/s
Mistral 7B v0.3 7.3B Runs (tight) Q3 5.2 GB ~20 tok/s
DeepSeek-R1 671B-A37B (MoE)MoE 671B Won't fit — 303 GB —
Qwen3 235B-A22B (MoE)MoE 235B Won't fit — 98 GB —
Llama 4 Scout (MoE)MoE 109B Won't fit — 46 GB —
Qwen2.5 72B 72.7B Won't fit — 33 GB —
Llama 3.3 70B 70.6B Won't fit — 32 GB —
DeepSeek-R1 Distill Llama 70B 70.6B Won't fit — 32 GB —
Mixtral 8x7B (MoE)MoE 46.7B Won't fit — 21 GB —
Qwen3 32B 32.8B Won't fit — 16 GB —
Qwen2.5 32B 32.8B Won't fit — 16 GB —
Qwen2.5-Coder 32B 32.8B Won't fit — 16 GB —
DeepSeek-R1 Distill Qwen 32B 32.8B Won't fit — 16 GB —
Qwen3 30B-A3B (MoE)MoE 30.5B Won't fit — 14 GB —
Gemma 3 27B 27.4B Won't fit — 16 GB —
Gemma 2 27B 27.2B Won't fit — 15 GB —
Mistral Small 3 24B 23.6B Won't fit — 12 GB —
Qwen3 14B 14.8B Won't fit — 8.0 GB —
Qwen2.5 14B 14.8B Won't fit — 8.3 GB —
DeepSeek-R1 Distill Qwen 14B 14.8B Won't fit — 8.3 GB —
Phi-4 14B 14.7B Won't fit — 8.3 GB —
Mistral Nemo 12B 12.2B Won't fit — 6.9 GB —
Gemma 3 12B 12.2B Won't fit — 8.7 GB —
Gemma 2 9B 9.2B Won't fit — 7.1 GB —

Similar hardware

FAQ

What is the best LLM for the Mac · M1/M2/M3 (base), 8 GB?

For general use, Gemma 3 4B is the strongest model that runs well on the Mac · M1/M2/M3 (base), 8 GB. See the picks-by-use-case below for coding, reasoning and more.

How much can the Mac · M1/M2/M3 (base), 8 GB run?

The Mac · M1/M2/M3 (base), 8 GB has 8 GB of unified memory, of which about 5.8 GB is usable for a model. That runs 5 of our tracked models well and 3 more at a tight quantisation.

Estimates at 8K context — see how we compute these.