Skip to content

Unified Memory Bandwidth and Tokens Per Second, Derived

10 min read · updated August 11, 2026

Nobody publishes tokens per second for your model on your Mac, and no figure on this page was measured. What can be derived, from two numbers that are both published, is a ceiling that no amount of software improvement will pass — and a ceiling is a more useful thing to reason with than a benchmark from somebody else’s machine.

The formula, and what it assumes

Generating one token means one forward pass, and one forward pass at batch size one reads essentially every weight the model uses, once, from memory. The arithmetic per weight is two floating-point operations — a multiply and an add — which is far too little work to hide the time spent fetching it. So the step time is set by how fast bytes arrive:

tokens per second  <=  memory bandwidth (bytes/s)
                       -----------------------------
                       bytes read per token

Four assumptions, stated rather than buried:

  • Batch size one. One sequence, one user. Two sequences decoded together read the weights once for both, so the ceiling per sequence halves while total throughput stays put. Batching changes this whole page.
  • A dense model. Every weight participates. A mixture-of-experts model reads only its routed experts per token, so the denominator is much smaller than its file size — treated separately below.
  • Peak bandwidth is achievable. It is not. The published GB/s is what the memory system can deliver under ideal access patterns; a real kernel achieves some fraction of it, and that fraction is a property of the implementation, not of the chip.
  • Decode only. This says nothing about prefill, which is a batched matrix multiply over the whole prompt at once and is limited by arithmetic rather than by bandwidth.

Because of the third assumption, everything below is an upper bound and the word “ceiling” is meant literally. A real number below it is expected; a real number above it means one of the four assumptions does not hold — most often batching or speculative decoding.

Bytes per token: the model

Use exact file sizes, not parameter counts times a guessed bytes-per-weight. MLX’s default affine quantization at 4 bits with group size 64 stores 4 bits per weight plus an fp16 scale and an fp16 bias per group, which is 4.5 bits per weight, and that prediction matches three real repositories read through the Hugging Face model API on 11 August 2026:

repository                                       safetensors bytes   params        bits/weight
mlx-community/Llama-3.1-8B-Instruct-4bit          4,517,489,037    8,030,261,248     4.50
mlx-community/Qwen3-14B-4bit                      8,307,898,514   14,768,307,200     4.50
mlx-community/Llama-3.3-70B-Instruct-4bit        39,688,567,605   70,553,706,496     4.50

The parameter counts are Hugging Face’s own safetensors index figures for meta-llama/Llama-3.3-70B-Instruct and the corresponding upstream repositories. So for a dense 4-bit MLX model, bytes per token from weights is the directory size: about 4.5 GB at 8B, 8.3 GB at 14B, 39.7 GB at 70B.

For a mixture-of-experts model this is wrong and the error is large. The 4-bit conversion of Qwen3-30B-A3B — 30,532,122,624 total parameters, of which roughly 3.3B are active per token — occupies 17.18 GB on disk, but only the routed experts plus the shared attention weights are read for any one token. Scaling by the active fraction puts the read at roughly 1.9 GB per token, an order of magnitude below the file size. All 17.18 GB must still be resident.

Bytes per token: the cache

The weights are not the only thing read. Attention reads the entire KV cache at every step, so the denominator grows with context length. The cache size per token is fixed by the architecture, and config.json gives you every term:

bytes per token = 2 (K and V) x layers x kv_heads x head_dim x bytes_per_value

Llama-3.3-70B: 80 layers, 8 kv_heads, head_dim 128, fp16
  = 2 x 80 x 8 x 128 x 2 = 327,680 bytes = 0.328 MB per token

So the extra read at a given context length is:

  4,096 tokens ->  1.34 GB      per decode step, on top of 39.7 GB of weights
  8,192 tokens ->  2.68 GB
 32,768 tokens -> 10.74 GB
131,072 tokens -> 42.95 GB      (the model's stated maximum position count)

At short context the cache is noise. At 32k it adds a quarter to the bytes read per token, and by the architecture’s maximum context the cache traffic exceeds the weight traffic. This is the mechanism behind “it got slower as the conversation went on”, and it is arithmetic rather than degradation. Quantizing the cache with kv_bits attacks the same term directly. The general behaviour of the cache is covered in the KV cache.

Worked ceilings

Apple publishes memory bandwidth per chip on its technical specification pages and in its Newsroom announcements. Reading those on 11 August 2026: M4 at 120GB/s and M4 Pro at 273GB/s on Apple’s Mac mini specifications; M4 Max at 410GB/s in its base configuration and 546GB/s with the 16-core CPU and 40-core GPU, and M3 Ultra at 819GB/s, on Apple’s Mac Studio specifications.

Dividing published bandwidth by the exact byte counts above, at short context so the cache term is negligible:

                    8B @ 4-bit   14B @ 4-bit   70B @ 4-bit
                       4.52 GB       8.31 GB      39.69 GB
M4        120 GB/s      26.5           14.4         does not fit
M4 Pro    273 GB/s      60.4           32.9         does not fit
M4 Max    546 GB/s     120.8           65.7          13.8
M3 Ultra  819 GB/s     181.2           98.6          20.6

ceiling in tokens/second, batch size one, dense, short context,
assuming 100% of published bandwidth — which no real kernel achieves

Read the last line before quoting any cell. These are not speeds. They are the number a perfect implementation could not exceed, and their practical use is comparative and diagnostic: they tell you that a 70B on an M3 Ultra is in the tens of tokens per second and not the hundreds, and that if you are seeing a third of the ceiling the problem is worth investigating while at ninety per cent of it there is nothing left to win.

Measuring the real number

The actual rate on your machine takes one command, and mlx-lm prints it without any instrumentation on your part:

mlx_lm.generate \
  --model mlx-community/Qwen3-14B-4bit \
  --prompt "Explain memory bandwidth in three sentences." \
  --max-tokens 300

The last three lines of output are the measurement:

Prompt: 14 tokens, 245.310 tokens-per-sec
Generation: 300 tokens, 41.882 tokens-per-sec
Peak memory: 8.734 GB

Generation is the decode rate this page bounds; Prompt is prefill and is a different, compute-bound number, usually an order of magnitude higher. Run it once at a short prompt and once at several thousand tokens of prompt, and the decode rate will fall between them by roughly the amount the KV term above predicts.

Three runs, discarding the first — the first pays for reading the weights off disk — gives a figure you can trust more than anything you will read on a forum, because it is your model, your quantization, your context length and your version of the runtime.