How we estimate

Every number on Intecca is computed from a small, transparent model of how large-language-model inference uses memory. We'd rather show our working than hand you a black box — so here's exactly what we calculate, and where it can be wrong.

Memory = weights + KV cache + overhead

A model needs to fit three things in your GPU VRAM, Apple unified memory, or system RAM:

1. Weights

The weights dominate. Their size is simply the parameter count times the bits-per-weight of the quantisation, divided by eight:

weights (bytes) = parameters × bits_per_weight ÷ 8

We use effective bits-per-weight taken from real GGUF file sizes, which include K-quant overhead — e.g. Q4_K_M ≈ 4.83 bpw, Q5_K_M ≈ 5.67, Q6_K ≈ 6.56, Q8_0 ≈ 8.5, and FP16 = 16. At Q4 that works out to roughly 0.6 GB per billion parameters.

2. KV cache

The attention cache grows with context length. For a transformer with grouped-query attention:

KV (bytes) = 2 × layers × kv_heads × head_dim × context × 2

The leading 2 counts keys and values; the trailing 2 is an FP16 cache (two bytes per element). Models with few KV heads (most modern ones use grouped-query attention) have far smaller caches. We compute this per model from its architecture; tables on the site assume an 8K-token context unless you change it in the calculator.

3. Overhead

Runtimes reserve memory for activations, CUDA/Metal context and compute scratch buffers. We add a practical floor of 0.75 GB + 4% of the weights. Real overhead varies by runtime and batch size.

How much memory you actually have

  • Dedicated GPU: VRAM minus ~0.6 GB for the driver and display.
  • Apple silicon: ~72% of unified memory, the rough default macOS lets the GPU wire. You can raise this with sudo sysctl iogpu.wired_limit_mb.
  • CPU / system RAM: total minus ~3 GB for the OS. Note CPU inference is slow regardless of how much fits.

We pick the highest-quality quantisation that still fits your usable memory. If a good quant (Q4 or better) fits, we call it Runs well; if only Q3/Q2 fits, Runs (tight); if nothing fits, Won't fit.

Speed (rough)

Token generation is memory-bandwidth-bound: each new token reads the active weights once. So our estimate is bandwidth × efficiency ÷ active-weight-bytes, with efficiency ≈ 0.72 of peak (and a further penalty for mixture-of-experts routing). This predicts the right ballpark and tier, not an exact figure — prompt processing, batching and your runtime all move it.

Where this is wrong

  • It's an estimate, not a guarantee. Treat the verdicts as "almost certainly / probably / no".
  • Architecture fields (layer and head counts) are community-sourced from model cards; parameters are authoritative, so the dominant weights term is accurate even where the KV term is approximate.
  • Speed estimates ignore prompt-processing time and assume a well-optimised runtime.
  • MoE models hold all experts in memory (so they're sized on total parameters) but only compute the active ones (so they're fast for their size).
  • We don't yet model speculative decoding, KV-cache quantisation, or weight offloading to RAM — all of which change the picture.

Freshness

The model list is versioned — currently 2026-06-22. This space moves weekly, so we add new releases as they land and revise figures when better numbers are known. If something looks off, it probably is worth a second look — tell us.