Prompt Caching Savings Calculator
Cacheable prefix, hit rate and three separate rates — ordinary, cache read and cache write — turned into a monthly saving and a break-even hit rate.
The prompt, split into the part that repeats and the part that does not
Your three input rates
Caching starts paying for itself above a 21.7% hit rate. You are at 70.0%.
- Cache hits per month
- 70,000
- Cache misses per month
- 30,000
- Prefix tokens billed per month
- 800,000,000
- Prefix cost with no caching
- $2,400
- — cache writes on misses
- $900.00
- — cache reads on hits
- $168.00
- Prefix cost with caching
- $1,068
- Saving
- $1,332
- Saving as a share of the prefix bill
- 55.5%
- Blended prefix rate
- $1.34 / 1M
- Break-even hit rate
- 21.7%
- Unchanged either way (rest of the input)
- $120.00
Prompt caching is usually presented as a discount, and it is not. It is a trade: you pay a premium the first time a prefix is seen in exchange for a large discount every time it is seen again. That makes it a break-even problem, and the break-even has a closed form. Caching pays when the blended rate — the write rate on misses, the read rate on hits — falls below what you would otherwise pay per input token, which happens above a hit rate of (write − input) ÷ (write − read). With a modest write premium that threshold is low, often well under a third, which is why caching is worth turning on for almost any repeated prefix. It is not why it is worth restructuring your prompt for.
The restructuring question is separate and more interesting. Caches match on prefixes, so the saving depends entirely on how much of your prompt is byte-identical from one request to the next, from the very first token. A system prompt that interpolates the current timestamp or the user's name at the top caches nothing at all, however static the following eight thousand tokens are. Moving the volatile parts to the end of the prompt is frequently the single highest-leverage change available, and it costs an afternoon.
What this does not model: the cost of a cache entry sitting unused, which is zero on per-token pricing but not on every provider; the quality effect of reordering a prompt, which is real and occasionally negative; and the interaction with batching, where an asynchronous window can push requests far enough apart that the entries expire. Measure your actual hit rate from usage logs rather than assuming it.