Skip to content

Running Out of Memory on a Mac Running a Local Model

10 min read · updated August 11, 2026

Three different things on a Mac present as “out of memory” when you run a local model, they print different messages, and the fix for one does nothing for the others. Start from the string you were given.

The error strings

From MLX’s Metal allocator, verbatim:

[metal::malloc] Attempting to allocate 12884901888 bytes which is greater
than the maximum allowed buffer size of 11453246976 bytes.
[metal::malloc] Resource limit (499000) exceeded.
[malloc] Unable to allocate 4294967296 bytes.

In Python these surface as a RuntimeError with that text. From llama.cpp’s Metal backend, the two you will see instead:

ggml_metal_buffer_init: error: failed to allocate buffer, size = 12288.00 MiB
ggml_metal_synchronize: error: command buffer 0 failed with status 5

And the fourth case has no error string at all: nothing fails, the model answers, and each token takes several seconds while the fans run and the rest of the machine stops responding. That is the swap case, and it is the one this page’s keyword is really about.

These strings are copied from the current source of both projects. Both rename functions and reword messages between releases, so match on the distinctive fragment — metal::malloc, failed to allocate buffer, command buffer — rather than on the whole line.

Which failure you have

  • greater than the maximum allowed buffer size is not about total memory. Metal caps the size of a single allocation, and one tensor exceeded that cap while plenty of memory remained free. It is reported by MLX as max_buffer_length. This is a shape problem, not a capacity problem: an enormous batch, a very long single sequence, or a model with an unusually large individual tensor.
  • Resource limit is a count of live Metal resources rather than bytes, and it usually means a program is creating buffers in a loop without letting them go.
  • Unable to allocate and llama.cpp’s failed to allocate buffer are the honest capacity failure. The request was larger than what the system was willing to give the GPU. This is the working-set ceiling, not your installed RAM.
  • command buffer 0 failed with status 5 is a Metal command buffer that ended in an error state — status 5 is MTLCommandBufferStatusError. Under memory stress it is a downstream symptom of one of the above; llama.cpp then logs the localised Metal description on the following line, and that line is the one worth reading.
  • No error, just crawling. Everything was allocated successfully, and the operating system is now servicing part of it from the SSD.

Reading memory pressure and swap

Open Activity Monitor, go to the Memory tab, and read three things rather than the top-line number.

  • The Memory Pressure graph. Green means the system is satisfying demand. Yellow means it is compressing pages to keep up. Red means it is paging to disk. Yellow during a model load and settling back to green is normal; sustained red during generation is the failure.
  • Swap Used. A non-zero value on macOS is not itself alarming — a long-running system accumulates some. A value that climbs while a model generates is the diagnosis, because it means the weights themselves are being paged.
  • Compressed. A large compressed figure is the stage before swap, and it is where the CPU cost starts: the system is spending real time compressing and decompressing pages to avoid the disk.

The reason this is so much worse than a clean failure is arithmetic. Decoding reads the whole model once per token; if part of it lives on the SSD, that part is fetched at storage speed rather than at the hundreds of gigabytes per second the memory delivers. Every token pays it, so the rate does not degrade gracefully — it collapses. This is the same bandwidth argument as the unified memory ceiling, with a much slower device substituted into the denominator.

Two commands are worth having:

sysctl vm.swapusage
memory_pressure

The first prints total, used and free swap. The second prints the system-wide pressure figure, which is the same signal the graph draws and is easier to watch in a terminal while a generation runs.

Fixes, matched to causes

  1. Find the real ceiling before changing anything. Run python -c "import mlx.core as mx; print(mx.device_info())" and compare max_recommended_working_set_size against your model’s size on disk. If the model is larger, nothing below fixes it and you need step 5.
  2. If the KV cache is the growth, cap or quantize it. Add --max-kv-size 8192 to bound it, or --kv-bits 4 to quarter it. For a 70B-shaped model the cache is 0.328 MB per token, so a 32k context is 10.7 GB — frequently the entire difference between fitting and not.
  3. Reclaim what MLX is holding. In a long-lived process, mx.clear_cache() releases the buffer cache, and mx.set_cache_limit(0) stops it accumulating at all — at some allocation cost. Drop any prompt caches you are keeping; they are real memory and are not the buffer cache.
  4. Free memory the obvious way. A browser with many tabs is routinely several gigabytes, and on a machine where the model needs eighty per cent of the pool, closing it is not a trivial gesture. Check what else is resident in Activity Monitor’s memory column before assuming the model is at fault.
  5. Reduce the model. The lever that always works. 4-bit MLX is 0.5625 bytes per parameter and 3-bit is 0.4375, so a step down is about a twenty-two per cent reduction in weights; a step down in parameter count is much larger. Choosing a quantization level for your memory budget works out which step you need rather than guessing.
  6. Only then, raise the wired limit. sudo sysctl iogpu.wired_limit_mb=<size_in_megabytes> lets the GPU hold more, and MLX’s mx.set_wired_limit() — macOS 15.0 or newer, and strictly less than total memory — applies it within a process. This is last because it takes headroom from the operating system: too aggressive and the machine becomes unresponsive rather than the model failing cleanly, and the setting does not survive a reboot.

For the buffer-size error specifically, none of the above applies. Reduce the batch size or the single-sequence length that produced the oversized tensor; the ceiling is per allocation and is not raised by having more memory free.

Making it not happen again

The durable fix is to size the model against the machine’s reported working set rather than against its headline RAM, and to include the KV cache at the context length you actually use. Both are one arithmetic step, and the method is in how much RAM a 70B needs on a Mac.

If a model is going to run unattended — a server, a scheduled job — set --max-kv-size explicitly rather than leaving the cache unbounded. An unbounded cache means a single long input can push a process that has been stable for weeks into swap, and the failure will present as a machine that stopped responding rather than as a job that errored.

The allocator controls, the wired limit and the working-set query are all covered properly in how MLX manages memory on Apple Silicon.