Apple

Best local LLMs for the Mac Studio · M3 Ultra, 512 GB

512 GB unified memory · 819 GB/s · ~369 GB usable. Enough unified memory to load DeepSeek-R1 671B at Q4 in one machine.

Try it with your context & use case

Preset to the Mac Studio · M3 Ultra, 512 GB. Change the context length or filter by use case.

29 models run well and 1 run tight on Mac Studio · M3 Ultra, 512 GB (any context — just fitting the weights) — 369 GB usable (≈72% of unified memory).

Sort:
ModelSizeFitBest quantNeedsMemorySpeed
Qwen3 235B-A22B (MoE)MoE235BRuns wellQ8_0243 GB
~16 tok/s
Llama 4 Scout (MoE)MoE109BRuns wellFP16212 GB
~11 tok/s
Qwen2.5 72B72.7BRuns wellFP16142 GB
~4.1 tok/s
Llama 3.3 70B70.6BRuns wellFP16138 GB
~4.2 tok/s
DeepSeek-R1 Distill Llama 70B70.6BRuns wellFP16138 GB
~4.2 tok/s
Mixtral 8x7B (MoE)MoE46.7BRuns wellFP1691 GB
~15 tok/s
Qwen3 32B32.8BRuns wellFP1664 GB
~9.0 tok/s
Qwen2.5 32B32.8BRuns wellFP1664 GB
~9.0 tok/s
Qwen2.5-Coder 32B32.8BRuns wellFP1664 GB
~9.0 tok/s
DeepSeek-R1 Distill Qwen 32B32.8BRuns wellFP1664 GB
~9.0 tok/s
Qwen3 30B-A3B (MoE)MoE30.5BRuns wellFP1660 GB
~58 tok/s
Gemma 3 27B27.4BRuns wellFP1654 GB
~11 tok/s
Gemma 2 27B27.2BRuns wellFP1653 GB
~11 tok/s
Mistral Small 3 24B23.6BRuns wellFP1646 GB
~12 tok/s
Qwen3 14B14.8BRuns wellFP1629 GB
~20 tok/s
Qwen2.5 14B14.8BRuns wellFP1629 GB
~20 tok/s
DeepSeek-R1 Distill Qwen 14B14.8BRuns wellFP1629 GB
~20 tok/s
Phi-4 14B14.7BRuns wellFP1629 GB
~20 tok/s
Mistral Nemo 12B12.2BRuns wellFP1624 GB
~24 tok/s
Gemma 3 12B12.2BRuns wellFP1624 GB
~24 tok/s
Gemma 2 9B9.2BRuns wellFP1619 GB
~32 tok/s
Qwen3 8B8.2BRuns wellFP1617 GB
~36 tok/s
Llama 3.1 8B8.0BRuns wellFP1616 GB
~37 tok/s
Qwen2.5-Coder 7B7.6BRuns wellFP1615 GB
~39 tok/s
Mistral 7B v0.37.3BRuns wellFP1615 GB
~41 tok/s
Gemma 3 4B4.3BRuns wellFP169.1 GB
~69 tok/s
Qwen3 4B4BRuns wellFP168.5 GB
~74 tok/s
Llama 3.2 3B3.2BRuns wellFP167.0 GB
~92 tok/s
Llama 3.2 1B1.2BRuns wellFP163.2 GB
~238 tok/s
DeepSeek-R1 671B-A37B (MoE)MoE671BRuns (tight)Q3_K_M318 GB
~21 tok/s

Estimates, computed in your browser — VRAM, quantisation and speed vary with your runtime and settings. How we estimate →

Top picks for the Mac Studio · M3 Ultra, 512 GB

Best for general chat & assistance

Qwen3 235B-A22B (MoE)

235B · Q8 · ~16 tok/s

Best for coding

Qwen3 235B-A22B (MoE)

235B · Q8 · ~16 tok/s

Best for reasoning & math

Qwen3 235B-A22B (MoE)

235B · Q8 · ~16 tok/s

Best for writing

Llama 3.3 70B

70.6B · FP16 · ~4.2 tok/s

Best for vision (image input)

Llama 4 Scout (MoE)

109B · FP16 · ~11 tok/s

Best for low-end & edge hardware

Gemma 3 4B

4.3B · FP16 · ~69 tok/s

Best for tool use & agents

Qwen3 235B-A22B (MoE)

235B · Q8 · ~16 tok/s

Every model on the Mac Studio · M3 Ultra, 512 GB

ModelSizeFitBest quantNeedsSpeed
Qwen3 235B-A22B (MoE)MoE 235B Runs well Q8 244 GB ~16 tok/s
Llama 4 Scout (MoE)MoE 109B Runs well FP16 213 GB ~11 tok/s
Qwen2.5 72B 72.7B Runs well FP16 144 GB ~4.1 tok/s
Llama 3.3 70B 70.6B Runs well FP16 140 GB ~4.2 tok/s
DeepSeek-R1 Distill Llama 70B 70.6B Runs well FP16 140 GB ~4.2 tok/s
Mixtral 8x7B (MoE)MoE 46.7B Runs well FP16 92 GB ~15 tok/s
Qwen3 32B 32.8B Runs well FP16 66 GB ~9.0 tok/s
Qwen2.5 32B 32.8B Runs well FP16 66 GB ~9.0 tok/s
Qwen2.5-Coder 32B 32.8B Runs well FP16 66 GB ~9.0 tok/s
DeepSeek-R1 Distill Qwen 32B 32.8B Runs well FP16 66 GB ~9.0 tok/s
Qwen3 30B-A3B (MoE)MoE 30.5B Runs well FP16 61 GB ~58 tok/s
Gemma 3 27B 27.4B Runs well FP16 58 GB ~11 tok/s
Gemma 2 27B 27.2B Runs well FP16 56 GB ~11 tok/s
Mistral Small 3 24B 23.6B Runs well FP16 48 GB ~12 tok/s
Qwen3 14B 14.8B Runs well FP16 31 GB ~20 tok/s
Qwen2.5 14B 14.8B Runs well FP16 31 GB ~20 tok/s
DeepSeek-R1 Distill Qwen 14B 14.8B Runs well FP16 31 GB ~20 tok/s
Phi-4 14B 14.7B Runs well FP16 31 GB ~20 tok/s
Mistral Nemo 12B 12.2B Runs well FP16 26 GB ~24 tok/s
Gemma 3 12B 12.2B Runs well FP16 27 GB ~24 tok/s
Gemma 2 9B 9.2B Runs well FP16 21 GB ~32 tok/s
Qwen3 8B 8.2B Runs well FP16 18 GB ~36 tok/s
Llama 3.1 8B 8.0B Runs well FP16 17 GB ~37 tok/s
Qwen2.5-Coder 7B 7.6B Runs well FP16 16 GB ~39 tok/s
Mistral 7B v0.3 7.3B Runs well FP16 16 GB ~41 tok/s
Gemma 3 4B 4.3B Runs well FP16 10 GB ~69 tok/s
Qwen3 4B 4B Runs well FP16 9.6 GB ~74 tok/s
Llama 3.2 3B 3.2B Runs well FP16 7.8 GB ~92 tok/s
Llama 3.2 1B 1.2B Runs well FP16 3.4 GB ~238 tok/s
DeepSeek-R1 671B-A37B (MoE)MoE 671B Runs (tight) Q3 349 GB ~21 tok/s

Similar hardware

FAQ

What is the best LLM for the Mac Studio · M3 Ultra, 512 GB?

For general use, Qwen3 235B-A22B (MoE) is the strongest model that runs well on the Mac Studio · M3 Ultra, 512 GB. See the picks-by-use-case below for coding, reasoning and more.

How much can the Mac Studio · M3 Ultra, 512 GB run?

The Mac Studio · M3 Ultra, 512 GB has 512 GB of unified memory, of which about 369 GB is usable for a model. That runs 29 of our tracked models well and 1 more at a tight quantisation.

Estimates at 8K context — see how we compute these.