How we estimate
Every number on Intecca is computed from a small, transparent model of how large-language-model inference uses memory. We'd rather show our working than hand you a black box — so here's exactly what we calculate, and where it can be wrong.
Memory = weights + KV cache + overhead
A model needs to fit three things in your GPU VRAM, Apple unified memory, or system RAM:
1. Weights
The weights dominate. Their size is simply the parameter count times the bits-per-weight of the quantisation, divided by eight:
weights (bytes) = parameters × bits_per_weight ÷ 8
We use effective bits-per-weight taken from real GGUF file sizes, which include K-quant overhead — e.g. Q4_K_M ≈ 4.83 bpw, Q5_K_M ≈ 5.67, Q6_K ≈ 6.56, Q8_0 ≈ 8.5, and FP16 = 16. At Q4 that works out to roughly 0.6 GB per billion parameters.
2. KV cache
The attention cache grows with context length. For a transformer with grouped-query attention:
KV (bytes) = 2 × layers × kv_heads × head_dim × context × 2
The leading 2 counts keys and values; the trailing 2 is an FP16 cache (two bytes per element). Models with few KV heads (most modern ones use grouped-query attention) have far smaller caches. We compute this per model from its architecture; tables on the site assume an 8K-token context unless you change it in the calculator.
3. Overhead
Runtimes reserve memory for activations, CUDA/Metal context and compute scratch buffers. We add a
practical floor of 0.75 GB + 4% of the weights. Real overhead varies by runtime and
batch size.
How much memory you actually have
- Dedicated GPU: VRAM minus ~0.6 GB for the driver and display.
- Apple silicon: ~72% of unified memory, the rough default macOS lets the GPU wire. You can raise this with
sudo sysctl iogpu.wired_limit_mb. - CPU / system RAM: total minus ~3 GB for the OS. Note CPU inference is slow regardless of how much fits.
We pick the highest-quality quantisation that still fits your usable memory. If a good quant (Q4 or better) fits, we call it Runs well; if only Q3/Q2 fits, Runs (tight); if nothing fits, Won't fit.
Speed (rough)
Token generation is memory-bandwidth-bound: each new token reads the active weights once. So our
estimate is bandwidth × efficiency ÷ active-weight-bytes, with efficiency ≈ 0.72 of peak
(and a further penalty for mixture-of-experts routing). This predicts the right ballpark and tier, not
an exact figure — prompt processing, batching and your runtime all move it.
Where this is wrong
- It's an estimate, not a guarantee. Treat the verdicts as "almost certainly / probably / no".
- Architecture fields (layer and head counts) are community-sourced from model cards; parameters are authoritative, so the dominant weights term is accurate even where the KV term is approximate.
- Speed estimates ignore prompt-processing time and assume a well-optimised runtime.
- MoE models hold all experts in memory (so they're sized on total parameters) but only compute the active ones (so they're fast for their size).
- We don't yet model speculative decoding, KV-cache quantisation, or weight offloading to RAM — all of which change the picture.
Freshness
The model list is versioned — currently 2026-06-22. This space moves weekly, so we add new releases as they land and revise figures when better numbers are known. If something looks off, it probably is worth a second look — tell us.