What Happens When a Model Doesn't Fit in VRAM
10 min read · updated August 11, 2026
Two things can happen, and they look nothing alike. Either the allocation fails and you get an error, or it succeeds by putting part of the model somewhere much slower and says nothing. The second is the one that wastes an afternoon.
The error strings
If you searched for one of these, the cause is below it. From llama.cpp and its CUDA backend, a line of this shape:
ggml_backend_cuda_buffer_type_alloc_buffer: allocating 4832.00 MiB on device 0: cudaMalloc failed: out of memory llama_model_load: error loading model: unable to allocate CUDA0 buffer
From PyTorch and anything built on it:
torch.cuda.OutOfMemoryError: CUDA out of memory. Tried to allocate 2.00 GiB. GPU 0 has a total capacity of 23.99 GiB of which 512.00 MiB is free.
From Ollama, a message of the form:
Error: model requires more system memory (34.2 GiB) than is available (15.6 GiB)
All three say the same thing: a specific allocation of a stated size did not fit in the stated remaining space. The size in the message is the useful part — it tells you whether you are 200 MiB short or 20 GiB short, and those have different fixes.
cudaMalloc failed, out of memory or unable to allocate rather than on the whole line.Two different failures
- Hard failure. You asked the runtime to put everything on the GPU —
-ngl 99in llama.cpp, or a device map that pins every layer — and it could not. The process stops. This is the good outcome: it is loud, immediate and costs nothing but a restart. - Silent spill. The runtime is allowed to place some layers on the CPU. It fits what it can into VRAM, puts the rest in system RAM, and serves the model at a fraction of the speed with no warning beyond a line in the load log saying how many layers went to the GPU. Ollama does this automatically by default; llama.cpp does it whenever
-nglis below the layer count, including when you omitted the flag entirely.
A third case sits between them: the weights fit but the KV cache for the requested context does not, so the failure happens at load with a size that looks unrelated to the model. If the number in the error is suspiciously round and scales with your -c value, that is what you are looking at.
A fourth is worth naming because it is the most confusing: the model loads, serves short requests happily, and dies on a long one. That is the compute buffer rather than the weights or the cache. It is sized for the widest tensor a forward pass will evaluate, which depends on how many tokens are processed in a single batch during prefill, so a 10,000-token prompt allocates more scratch than a 100-token one. The allocation that fails names a size that is neither the model nor the cache, and the fix is a smaller physical batch rather than a smaller model.
Deriving the cliff
Single-stream decode reads every weight once per token, so time per token is bytes divided by the bandwidth those bytes come over. When the model is split, the two halves come over very different buses:
t_token = (f * B) / bw_vram + ((1 - f) * B) / bw_host
B total weight bytes
f fraction resident in VRAM
bw_vram GPU memory bandwidth, order 300-1000 GB/s on a modern card
bw_host system memory bandwidth, order 50-100 GB/s for dual-channel DDR5,
or PCIe bandwidth if the weights stream over the bus each passPut numbers in. Llama 3.1 8B at Q4_K_M is 4.58 GiB = 4.92e9 bytes. Take a card with 1,008 GB/s and a host path at 89.6 GB/s — dual-channel DDR5-5600, which is 5600 MT/s × 8 bytes × 2 channels:
f = 1.00 t = 4.92e9/1008e9 = 4.88 ms -> 205 tok/s ceiling f = 0.90 t = 4.43e9/1008e9 + 0.49e9/89.6e9 = 9.87 ms -> 101 tok/s (2.0x slower) f = 0.50 t = 2.46e9/1008e9 + 2.46e9/89.6e9 = 29.9 ms -> 33 tok/s (6.1x slower)
Ten per cent of the model off the card halves the rate. Half of it off the card costs a factor of six. That non-linearity is the whole reason a spill feels like a cliff rather than a slope: the slow term dominates the moment it exists at all, because the two bandwidths differ by more than a factor of ten.
These are ceilings computed from stated bandwidths, not measurements — no runtime reaches its roofline, and the real ratio depends on your memory configuration and on whether the host-side layers are computed on the CPU or streamed to the GPU. The shape is what transfers. The underlying argument is in why memory bandwidth sets inference speed.
Two things make the real cliff steeper than that model. Layers are the unit of offload, so the split is coarse — you cannot spill three per cent of a model, only whole layers, and on a 32-layer model the smallest step is about three per cent at a time. And the offloaded layers are usually computed on the CPU rather than streamed to the GPU, so the arithmetic for those layers also runs at CPU throughput, and the two halves cannot overlap because layers are sequential. The derivation above counts only memory traffic and is therefore an optimistic bound on a pessimistic situation.
Diagnosing which one you have
- Read the load log for the offload line. llama.cpp prints how many of the N layers were assigned to each device. If it says 28 of 33, five layers are on the CPU and you are in the slow case.
- Run
nvidia-smi --query-gpu=memory.used,memory.total,utilization.gpu --format=csv -l 1during generation. A spilled model shows VRAM used well below total — the runtime stopped filling it — while generation is slow. - Watch a CPU monitor. In a partial offload the CPU cores handling the resident-in-RAM layers run hot during generation. Full GPU residency leaves the CPU nearly idle.
- Compare tokens per second against the ceiling above: weight bytes divided by your card’s published bandwidth. Coming in at a tenth of that with no other explanation means part of the model is not where you think it is.
The fix
- Force the failure. Set
-ngl 99and an explicit-c. If it cannot all fit you get an error instead of a mystery, and the error names the size it could not allocate. - Shrink the cache before the weights. Lowering
-cor moving to--cache-type-k q8_0 --cache-type-v q8_0frees memory with less quality cost than another quant step down. The numbers are on the 8 GiB budget page. - Step down a quant, then a size class. Q4_K_M to Q4_K_S is 4.6% of the weights; Q5_K_M to Q4_K_M is 14%. Past that, a smaller model at a higher quant usually beats the same model mangled further.
- Accept the offload deliberately. If you need the larger model and can live with single-digit tokens per second, set
-nglto the largest value that loads and know what you bought. The trade is legitimate for batch work; it is not for anything a person waits on.