Quantization
Quantization stores a model’s weights at lower numeric precision to shrink it. A model trained in 16-bit (FP16) can be packed down to 8, 5, 4 or even 2 effective bits per weight — cutting memory roughly in proportion.
The common GGUF levels are labelled like Q4_K_M (≈4.8 bits/weight), Q5_K_M (≈5.7), Q6_K (≈6.6) and Q8_0 (≈8.5). Q4_K_M is the usual sweet spot: about a quarter the size of FP16 with quality loss most people can’t notice. Dropping to Q3 or Q2 saves more but degrades output, so it’s a last resort to fit a bigger model.
VRAM
VRAM is the fast memory on a graphics card. To run a model on a GPU, its weights, attention cache and a little overhead all have to fit in VRAM at once — if they don’t, the model either won’t load or spills to slow system RAM.
As a rule of thumb at Q4 you need about 0.6 GB of VRAM per billion parameters, plus room for context. That’s why VRAM, not raw compute, is usually the wall you hit first — and why a 24 GB card is the threshold for 32B-class models.
Unified memory (Apple silicon)
Apple-silicon Macs share one pool of memory between the CPU and GPU. That means a 64 GB Mac can devote far more memory to a model than a typical consumer graphics card — enough to run a 70B model that would otherwise need two GPUs.
The catch is bandwidth and a soft cap: macOS only lets the GPU “wire” roughly 72% of unified memory by default (raisable with a sysctl), and even the fastest Macs trail a high-end GPU on raw bandwidth, so tokens-per-second is lower than the memory size alone suggests.
KV cache
As a model reads your prompt and generates text, it caches the attention keys and values for every token so it doesn’t recompute them. This KV cache lives in memory alongside the weights and grows linearly with context length.
At long context it can add several gigabytes. Modern models reduce it with grouped-query attention (fewer KV heads), and some runtimes can quantize the cache itself to halve its size.
Context window
The context window is how many tokens (roughly ¾ of a word each) a model can consider at once — your prompt plus its reply. Bigger windows let you feed in long documents or codebases.
A long context isn’t free: it inflates the KV cache and slows the first response. Most local models advertise 32K–128K windows, but you’ll often run them at 8K–16K to save memory unless you actually need the length.
GGUF
GGUF is the file format used by llama.cpp and tools built on it (Ollama, LM Studio, Jan). One .gguf file bundles the quantized weights and the metadata needed to run them, which is why most local-model downloads come as GGUF.
The quant level is in the filename (e.g. model-Q4_K_M.gguf), and the file size on disk is close to the memory the weights will use — a handy sanity check against any calculator.
Grouped-query attention (GQA)
Grouped-query attention lets several attention “query” heads share a single key/value head. It barely affects quality but sharply reduces the KV cache, which is why almost every recent model uses it.
In practice GQA is the reason a modern 8B model needs far less context memory than an older one of the same size — fewer KV heads means a smaller cache per token.
Mixture of experts (MoE)
A mixture-of-experts model contains many “expert” sub-networks but only activates a few per token. Qwen3 30B-A3B, for example, holds 30B parameters but computes only ~3B at a time.
The consequence for hardware: MoE models need memory for all their parameters (you must load every expert) but run at the speed of just the active ones. That makes them a great fit for high-memory, modest-bandwidth machines like Macs.
Tokens per second
Tokens per second (tok/s) measures how fast a model generates text. Above ~15 tok/s reads as comfortably interactive; single digits feel sluggish; 40+ is fast.
Because generation reads the model’s weights from memory once per token, speed tracks memory bandwidth divided by the (active) model size — not the GPU’s raw compute. A smaller quant is both smaller and faster for this reason.
Want the formulas behind the numbers? See how we estimate, or just try the calculator.