Skip to content

Fitting a Model and Its Context Window in 24GB of VRAM

10 min read · updated August 11, 2026

Twenty-four gigabytes buys you a genuine choice rather than a constraint. The same card runs a 34B, a 13B at eight bits, or an 8B with more context than most hosted models offer — and the arithmetic says exactly what each costs.

The budget

NVIDIA publishes 24 GB GDDR6X on a 384-bit interface for the GeForce RTX 4090; the same capacity applies to the 3090 and to several workstation parts. Card memory is reported in binary units, so 24 GB here means 24 GiB = 25,769,803,776 bytes.

C = (V - W - R) / k       V = 24 GiB, R assumed 1.2 GiB

R is larger than on a small card, not smaller: the compute buffer scales
with the batch size the runtime is configured for, and a bigger card is
usually asked to do more per pass.

The reserve is worth arguing about at this size, because a 24 GiB card is usually the one doing serious work and the compute buffer is allocated for the widest tensor the graph will evaluate. Raising the physical batch size to speed up prompt processing makes that buffer larger, and on a configuration sized to the last few hundred megabytes that change is what breaks it. If a setup that loaded yesterday stops loading after a flag change, look at the batch settings before you look at the model.

Three regimes

Every configuration on this card is one of three trades, and the numbers below come from the same equation with different models substituted:

model                    quant    weights   free     context
Yi-34B                   Q4_K_S   18.68 GiB  4.12 GiB  17,980
Yi-34B                   Q4_K_M   19.59      3.21      14,006
CodeLlama-34B            Q4_K_M   19.23      3.57      19,513
Llama 2 13B              Q8_0     12.88      9.92      13,001
Llama 3.1 8B             F16      14.96      7.84      64,241
Llama 3.1 8B             Q8_0      7.95     14.85     121,676
Llama 3.1 8B  Q8_0, q8_0 cache     7.95     14.85     243,352

Regime one: the largest model that fits

A 34B at Q4. Yi-34B is 34,388,917,248 parameters over 60 layers with eight key/value heads, so 240 KiB of cache per token at fp16:

weights  34,388,917,248 * 4.8944 / 8 = 19.59 GiB
free     24.00 - 19.59 - 1.20        =  3.21 GiB
context  3.21 GiB / 240 KiB per token = 14,006 tokens

Fourteen thousand tokens, and the Q5_K_M row does not exist — 22.83 GiB of weights plus 1.2 GiB of runtime is 24.03, which is over the card before a single token of context. This is the regime with the least headroom on the whole card, and the one where an underestimated R turns a working setup into a load failure. The full derivation for this class is on the 34B page.

This regime is also the slowest of the three, and for a reason the memory table does not show. Single-stream generation reads every resident weight once per token, so time per token scales with the weight bytes: 19.59 GiB against 7.95 GiB for the 8B at Q8_0 is 2.5x the traffic and therefore roughly 2.5x the time per token, on identical hardware. Choosing the largest model that fits is choosing the slowest configuration the card supports, which is fine for considered answers and poor for anything interactive. The argument is developed on the batch-size-one page.

Regime two: precision instead of size

A 13B at Q8_0 is 12.88 GiB — 8.5 bits per weight, which is close enough to the original that the quantization is rarely the limiting factor in output quality. It leaves 9.92 GiB, but Llama 2 13B’s multi-head attention costs 800 KiB per token, so that is only 13,001 tokens.

The interesting version of this regime is an 8B at F16: 14.96 GiB of weights, no quantization at all, and 64,241 tokens of context. If you are evaluating a fine-tune, comparing against a quantized build, or doing anything where you need to be sure the weights are not the variable, this is the configuration a 24 GiB card exists for. The quality question is discussed under how good local models actually are.

The 13B row is a warning rather than a recommendation. Twelve point nine gigabytes of weights leaves 9.92 GiB free, which on a grouped-query model would be eighty thousand tokens and on this one is thirteen thousand. Nearly ten gigabytes of a 24 GiB card spent on thirteen thousand tokens is the clearest illustration in this cluster that the cache term, not the weights term, is what a big card buys you — and that it only buys it on an architecture that can use it.

Regime three: context instead of both

An 8B at Q8_0 is 7.95 GiB and leaves 14.85 GiB. At 128 KiB per token that is 121,676 tokens — effectively the model’s entire 128k window, on one consumer card, with the weights at eight bits. Move the cache to q8_0 and the arithmetic gives 243,352, which is past what the model was trained for and therefore not useful, but it means the cache is no longer the constraint at any context the model can handle.

This is the regime that a hosted API is worst at replacing, and the reason people buy this card. A long window on a local model means the document never leaves the machine, the cost of a 100,000-token prompt is electricity rather than per-token billing, and you can re-run the same enormous prompt fifty times while iterating on the instructions without watching a meter. Whether the model is good enough at that length is a different question from whether it fits, and the answer to the second one is now yes.

Fitting a 128k cache is not the same as the model using 128k well. Attention cost grows with the square of the sequence during prefill, so a full window is slow to fill even when it fits, and effective quality at long range is a separate question from memory.

Why a 70B is not on this list

Llama 3.1 70B is 70,553,706,496 parameters. At Q2_K — 3.1593 bits/weight, the smallest k-quant llama.cpp publishes — that is 25.9 GiB of weights, already over the card. The i-quants reach lower: IQ2_M at 2.9294 bits/weight is 24.0 GiB and IQ2_XXS at 2.3824 is 19.6 GiB. So a 70B on a single 24 GiB card means roughly two bits per weight and no realistic context budget.

It is tempting to treat this as a near miss and go looking for the configuration that just about works. It is not a near miss. The shortfall between 24 GiB and a usefully quantized 70B is about 18 GiB, which is not recoverable by shaving the context or the reserve — those are gigabytes and this is tens of them. The two real answers are a second card or a smaller model, and the arithmetic says which long before an evening has been spent on flags.

Set against regime one, that is a 70B mangled to two bits versus a 34B at a comfortable four and a half, in the same memory. The card counts for running a 70B properly are on the 70B page, and the cost of the alternative — keeping some layers in system RAM — is derived in what happens when a model does not fit.