Prompt Caching in the Claude API: cache_control Breakpoints and the TTL
9 min read · updated August 11, 2026
Prompt caching is one marker on one content block, and almost every question about it — why did it not hit, where do I put it, does changing this break it — is answered by the same sentence: the cache key is a prefix of the rendered prompt.
The one rule: it is a prefix match
A cache_control marker on a block says “everything from the start of the prompt up to and including this block is a cache entry”. The lookup is on the exact bytes of that prefix. Change one character anywhere inside it — a timestamp, a reordered JSON key, a tool added to the array — and the entry does not match and nothing after that point can be read from cache either.
The prompt is rendered in a fixed order: tools, then system, then messages. This single fact explains most surprises. A marker on the last system block caches the tools too, because they precede it. Adding a tool invalidates a system-prompt cache, because the tools are earlier in the prefix. And interpolating the current date into the system prompt invalidates the entire cache on every request, because the prefix changes each time.
Where the marker goes
cache_control is a property of a content block, not of the request. It can go on a system text block, on a tool definition, or on a message content block:
{
"model": "claude-opus-4-6",
"max_tokens": 1024,
"tools": [ … ],
"system": [
{
"type": "text",
"text": "<20,000 tokens of internal API reference>",
"cache_control": {"type": "ephemeral"}
}
],
"messages": [
{"role": "user", "content": "Which endpoint returns deploy history?"}
]
}ephemeral is the only type. The number of markers per request is capped — four at the time of writing — which is enough for the shape most applications need: one at the end of the tools, one at the end of the system prompt, and one or two moving through the conversation as it grows.
For a multi-turn conversation, the marker moves. Put it on the last content block of the most recently appended turn, and each request reuses the entire prior conversation as its prefix while writing a new entry that includes the turn just added.
A request that hits and one that misses
The difference is usually a single interpolated value. This misses, every time:
"system": [
{
"type": "text",
"text": "Current time: 2026-08-11T09:14:22Z\n<20,000 tokens of reference>",
"cache_control": {"type": "ephemeral"}
}
]The timestamp is at the front of the prefix, so every request produces a different key. You pay the cache-write premium on every call and never read a single entry. The same content, restructured, hits:
"system": [
{
"type": "text",
"text": "<20,000 tokens of reference>",
"cache_control": {"type": "ephemeral"} // stable — this is the cache entry
},
{
"type": "text",
"text": "Current time: 2026-08-11T09:14:22Z" // volatile — after the breakpoint
}
]Nothing was removed. The volatile content moved after the breakpoint, so it is re-read on every request — 20 tokens instead of 20,000 — and the expensive part is served from cache. The general form of the fix is always this: sort content by how often it changes, stable first.
TTL, minimums and the economics
- Two TTLs. The default entry lives five minutes, and the timer resets on every read — a conversation with a turn every couple of minutes keeps its cache alive indefinitely. A one-hour TTL is available by adding
"ttl": "1h"to thecache_controlobject. - There is a minimum length. A prefix below the model’s minimum cacheable size is silently not cached — no error, just a
cache_creation_input_tokensof zero. The minimum differs per model and is not monotonic across generations, so a prompt that caches on one model may not on another. - Writes cost more than normal input. Writing an entry is charged at a premium over base input rates and reading it at a fraction of them. The five-minute TTL breaks even at roughly two requests against the same prefix; the one-hour TTL costs more to write and needs more reads to pay for itself, which makes it a bursty-traffic instrument rather than a default.
The invalidators you cannot see
The timestamp example above is the visible case. The expensive ones are the changes you did not write, because the prefix is compared as serialised bytes and several perfectly reasonable pieces of code produce different bytes for the same logical content.
- Non-deterministic serialisation. A tool schema built from a dictionary and serialised without sorted keys can render its properties in a different order between processes. The content is identical, the bytes are not, and the cache misses on every machine except the one you tested on. Sort keys when you serialise anything that lands in the prefix.
- Per-user or per-request tool sets. Filtering the
toolsarray by the caller’s permissions gives every permission combination its own cache entry, andtoolsrender first, so the divergence invalidates the system prompt behind it too. Sending a stable superset and enforcing permissions in your tool handler usually costs fewer tokens than the cache misses it prevents. - Anything generated per request. A UUID for tracing, a request counter, a randomly selected few-shot example, a “you are talking to Alex” line at the top of a shared system prompt. All of them are cheap to move after the breakpoint and ruinous in front of it.
- Silent library changes. An SDK that adds a field to the tool definitions it sends, or normalises whitespace differently after an upgrade, changes the prefix without a line of your code changing. If cache hit rate drops on a deploy that touched no prompts, this is the first thing to check.
There is also a shape that produces a cache which technically works and never pays for itself: writing an entry that is read zero or one times. A write costs more than an uncached read, so caching a prefix that is unique to a single request is strictly worse than not caching it. This is why caching per-user document context is usually a losing trade unless that user is coming back inside the TTL, and why the feature belongs on the parts of the prompt that are shared across callers.
Verifying with the usage fields
Do not reason about whether caching is working. The response tells you, in three fields on usage:
"usage": {
"input_tokens": 42, // read at full price this request
"cache_creation_input_tokens": 0, // written to cache this request
"cache_read_input_tokens": 20134 // served from cache this request
}The first request against a new prefix shows a large cache_creation_input_tokens and a zero read. The second shows the mirror image. If cache_read_input_tokens stays at zero across repeated requests you believe are identical, something in the prefix is not identical — diff the serialised request bodies rather than guessing, and look first at anything generated per request.
One more thing worth knowing: input_tokens is only the uncached remainder. Total prompt size is the sum of all three fields, so an agent whose input_tokens reads 42 after an hour of work has not been sending a 42-token prompt.