Skip to content

Fitting a Model and Its Context Window in 8GB of VRAM

9 min read · updated August 11, 2026

An 8 GiB card is the most common thing people try to run a local model on and the tightest budget where a useful one still fits. The question is not whether a 7B loads — it does — but how much conversation is left afterwards.

The budget equation

Rearrange the memory formula to solve for context rather than for total:

C = (V - W - R) / k

V   card capacity, in bytes. An "8GB" card is 8 GiB = 8,589,934,592 bytes.
W   weights = P * bpw / 8
R   runtime reserve: CUDA context, compute buffer, display framebuffer
k   kv bytes per token = 2 * L * H_kv * D * b

Every quantity on the right is knowable. W is derived on the 7B weights page, k from the model’s layer and head configuration, and R is the only one you have to measure. It is also the one every published table quietly sets to zero, which is why their numbers are optimistic.

What to reserve, and why it is not zero

Three things occupy the card before a single weight arrives, and a fourth appears at load time:

  • The CUDA primary context, created the first time any CUDA call runs in a process. It holds driver state and the loaded kernel images.
  • The display. If the card drives a monitor, the framebuffer and the desktop compositor are resident. On Windows this is typically the largest single item outside the model.
  • Other processes. A browser with hardware acceleration on is using your inference card.
  • The runtime’s compute buffer, allocated when the model loads. It holds activations for a forward pass and scales with the batch size and the logical batch size the runtime is configured for, not with total context.

A 1.0 GiB reserve is used below. It is an assumption, chosen to be conservative on a headless Linux box and roughly right on a Windows desktop that is not doing anything else. Replace it with the number your machine actually shows; the method is in how much VRAM CUDA reserves before your model loads.

Context per quant

Llama 3.1 8B (8.03B parameters, 32 layers, 8 KV heads, head dimension 128, so 128 KiB per token at fp16) and Mistral 7B (7.24B, same cache shape). Weights come from the bits-per-weight figures published in llama.cpp’s quantize README, which measures them on this exact 8B. With V = 8 GiB and R = 1.0 GiB:

Llama 3.1 8B          weights   free    max context (fp16 cache)
  Q3_K_M               3.74     3.26 GiB      26,741 tokens
  Q4_K_S               4.36     2.64          21,601
  Q4_K_M               4.58     2.42          19,861
  Q5_K_M               5.33     1.67          13,664
  Q6_K                 6.14     0.86           7,080
  Q8_0                 7.95    negative       will not load

Mistral 7B
  Q4_K_M               4.13     2.87          23,542
  Q5_K_M               4.81     2.19          17,953
  Q6_K                 5.53     1.47          12,016

Llama 3.2 3B
  Q4_K_M               1.83     5.17          48,397
  Q8_0                 3.18     3.82          35,769

The headline: an 8 GiB card runs an 8B at Q4 with about 20,000 tokens, or at Q6 with about 7,000. It does not run one at Q8 at all. And a 3B at Q8_0 — full eight-bit weights, essentially no quantization damage — leaves 35,000 tokens, which for many tasks is a better machine than an 8B squeezed to Q4 with a short window.

Read the free column rather than the weights column, because that is where the non-linearity lives. Going from Q4_K_M to Q5_K_M costs 0.75 GiB of weights, which sounds modest — but it is 31% of the free space, so the context falls by 31%, from 19,861 to 13,664. Another step to Q6_K costs 0.81 GiB and takes half of what is left. On a small card the marginal cost of a quant step is not the size of the step, it is the size of the step divided by the shrinking remainder, and it rises steeply. This is why 8 GiB configurations feel like they fall off a cliff at Q6 while a 24 GiB card barely notices the same change.

The same reasoning applies in the other direction and is the reason to consider going down a size class rather than down a quant. A 3B at Q4_K_M leaves 5.17 GiB free — more free space than an 8B has weights. Whether three billion parameters is enough is a question about your task and not about memory, but on this card the answer “a smaller model at full precision with a long window” is available and is rarely considered.

Quantizing the cache

The b term in the denominator is under your control. llama.cpp exposes separate types for the key and value caches, and moving both to q8_0 halves the per-token cost:

llama-server -m model.gguf -c 32768 \
  --cache-type-k q8_0 --cache-type-v q8_0 -ngl 99

Llama 3.1 8B at Q4_K_M on 8 GiB, R = 1.0:
  fp16 cache (2 B/elem)   128 KiB/token -> 19,861 tokens
  q8_0 cache (1 B/elem)    64 KiB/token -> 39,722 tokens

Twice the context for one byte per element instead of two. This is usually the highest-value lever on a small card, because the cache stores per-token activations rather than learned weights, and the errors do not compound across the whole model the way weight quantization errors do.

The asymmetry is worth understanding rather than taking on trust. A quantized weight is wrong for every token the model will ever generate, and the error is applied again at every step, so it is a permanent change to the function. A quantized cache entry is a stored activation for one particular token in one particular conversation; it is read by later tokens attending to it and then discarded when the session ends. There is no compounding across the model, only a slightly noisier view of what an earlier token contributed. The key cache is the more sensitive of the two, since keys determine which entries get attended to at all, which is why some builds let you set the two precisions separately and why dropping the value cache further is the safer half of that trade.

Flag names and defaults move between llama.cpp releases, and quantized value caches have historically required flash attention to be enabled on some backends. Check llama-server --help for the build you have rather than copying flags from a year-old post.

Reading it off the runtime

Do not trust the arithmetic over the log. llama.cpp prints its allocations at load time — the model buffer, the KV cache, and the compute buffer, per device — and those lines are the ground truth for your build, your model and your flags:

  1. Load with an explicit context: -c 8192 -ngl 99. Do not leave the context at the model default, which for a modern model may be 128k and will try to allocate a cache for all of it.
  2. Read the buffer sizes the loader prints and add them. Compare with nvidia-smi --query-gpu=memory.used --format=csv taken while the process is alive.
  3. The difference between the two is your R. Put that number back into the equation above and re-solve for the context you can afford.
  4. Raise -c until the load fails, then back off. The failure is immediate and harmless, which makes bisecting cheap.