NVIDIA · 2024-01

Best local LLMs for the RTX 4070 Super

12 GB VRAM · 504 GB/s · ~11 GB usable.

Check RTX 4070 Super price →

Try it with your context & use case

Preset to the RTX 4070 Super. Change the context length or filter by use case.

15 models run well and 1 run tight on RTX 4070 Super (any context — just fitting the weights) — 11 GB usable.

Sort:
ModelSizeFitBest quantNeedsMemorySpeed
Qwen3 14B14.8BRuns wellQ5_K_M11 GB
~35 tok/s
Qwen2.5 14B14.8BRuns wellQ5_K_M11 GB
~35 tok/s
DeepSeek-R1 Distill Qwen 14B14.8BRuns wellQ5_K_M11 GB
~35 tok/s
Phi-4 14B14.7BRuns wellQ5_K_M11 GB
~35 tok/s
Mistral Nemo 12B12.2BRuns wellQ6_K10 GB
~36 tok/s
Gemma 3 12B12.2BRuns wellQ6_K10 GB
~36 tok/s
Gemma 2 9B9.2BRuns wellQ8_010 GB
~37 tok/s
Qwen3 8B8.2BRuns wellQ8_09.2 GB
~42 tok/s
Llama 3.1 8B8.0BRuns wellQ8_09.0 GB
~43 tok/s
Qwen2.5-Coder 7B7.6BRuns wellQ8_08.6 GB
~45 tok/s
Mistral 7B v0.37.3BRuns wellQ8_08.2 GB
~47 tok/s
Gemma 3 4B4.3BRuns wellFP169.1 GB
~42 tok/s
Qwen3 4B4BRuns wellFP168.5 GB
~45 tok/s
Llama 3.2 3B3.2BRuns wellFP167.0 GB
~57 tok/s
Llama 3.2 1B1.2BRuns wellFP163.2 GB
~146 tok/s
Mistral Small 3 24B23.6BRuns (tight)Q2_K10 GB
~37 tok/s
DeepSeek-R1 671B-A37B (MoE)MoE671BWon't fit—273 GB
—
Qwen3 235B-A22B (MoE)MoE235BWon't fit—96 GB
—
Llama 4 Scout (MoE)MoE109BWon't fit—45 GB
—
Qwen2.5 72B72.7BWon't fit—30 GB
—
Llama 3.3 70B70.6BWon't fit—29 GB
—
DeepSeek-R1 Distill Llama 70B70.6BWon't fit—29 GB
—
Mixtral 8x7B (MoE)MoE46.7BWon't fit—20 GB
—
Qwen3 32B32.8BWon't fit—14 GB
—
Qwen2.5 32B32.8BWon't fit—14 GB
—
Qwen2.5-Coder 32B32.8BWon't fit—14 GB
—
DeepSeek-R1 Distill Qwen 32B32.8BWon't fit—14 GB
—
Qwen3 30B-A3B (MoE)MoE30.5BWon't fit—13 GB
—
Gemma 3 27B27.4BWon't fit—12 GB
—
Gemma 2 27B27.2BWon't fit—12 GB
—

Estimates, computed in your browser — VRAM, quantisation and speed vary with your runtime and settings. How we estimate →

Top picks for the RTX 4070 Super

Best for general chat & assistance

Qwen3 14B

14.8B · Q4 · ~41 tok/s

Best for coding

Qwen3 14B

14.8B · Q4 · ~41 tok/s

Best for reasoning & math

Qwen3 14B

14.8B · Q4 · ~41 tok/s

Best for writing

Qwen2.5 14B

14.8B · Q4 · ~41 tok/s

Best for vision (image input)

Gemma 3 12B

12.2B · Q4 · ~49 tok/s

Best for low-end & edge hardware

Gemma 3 4B

4.3B · FP16 · ~42 tok/s

Best for tool use & agents

Qwen3 14B

14.8B · Q4 · ~41 tok/s

Every model on the RTX 4070 Super

ModelSizeFitBest quantNeedsSpeed
Qwen3 14B 14.8B Runs well Q4 11 GB ~41 tok/s
Qwen2.5 14B 14.8B Runs well Q4 11 GB ~41 tok/s
DeepSeek-R1 Distill Qwen 14B 14.8B Runs well Q4 11 GB ~41 tok/s
Phi-4 14B 14.7B Runs well Q4 11 GB ~41 tok/s
Mistral Nemo 12B 12.2B Runs well Q5 10 GB ~42 tok/s
Gemma 3 12B 12.2B Runs well Q4 11 GB ~49 tok/s
Gemma 2 9B 9.2B Runs well Q6 11 GB ~48 tok/s
Qwen3 8B 8.2B Runs well Q8 10 GB ~42 tok/s
Llama 3.1 8B 8.0B Runs well Q8 10 GB ~43 tok/s
Qwen2.5-Coder 7B 7.6B Runs well Q8 9.0 GB ~45 tok/s
Mistral 7B v0.3 7.3B Runs well Q8 9.2 GB ~47 tok/s
Gemma 3 4B 4.3B Runs well FP16 10 GB ~42 tok/s
Qwen3 4B 4B Runs well FP16 9.6 GB ~45 tok/s
Llama 3.2 3B 3.2B Runs well FP16 7.8 GB ~57 tok/s
Llama 3.2 1B 1.2B Runs well FP16 3.4 GB ~146 tok/s
DeepSeek-R1 671B-A37B (MoE)MoE 671B Won't fit — 303 GB —
Qwen3 235B-A22B (MoE)MoE 235B Won't fit — 98 GB —
Llama 4 Scout (MoE)MoE 109B Won't fit — 46 GB —
Qwen2.5 72B 72.7B Won't fit — 33 GB —
Llama 3.3 70B 70.6B Won't fit — 32 GB —
DeepSeek-R1 Distill Llama 70B 70.6B Won't fit — 32 GB —
Mixtral 8x7B (MoE)MoE 46.7B Won't fit — 21 GB —
Qwen3 32B 32.8B Won't fit — 16 GB —
Qwen2.5 32B 32.8B Won't fit — 16 GB —
Qwen2.5-Coder 32B 32.8B Won't fit — 16 GB —
DeepSeek-R1 Distill Qwen 32B 32.8B Won't fit — 16 GB —
Qwen3 30B-A3B (MoE)MoE 30.5B Won't fit — 14 GB —
Gemma 3 27B 27.4B Won't fit — 16 GB —
Gemma 2 27B 27.2B Won't fit — 15 GB —
Mistral Small 3 24B 23.6B Won't fit — 12 GB —

Similar hardware

FAQ

What is the best LLM for the RTX 4070 Super?

For general use, Qwen3 14B is the strongest model that runs well on the RTX 4070 Super. See the picks-by-use-case below for coding, reasoning and more.

How much can the RTX 4070 Super run?

The RTX 4070 Super has 12 GB of VRAM, of which about 11 GB is usable for a model. That runs 15 of our tracked models well.

Estimates at 8K context — see how we compute these.