Skip to content

What Breaks in an Internal Chargeback Model After a Provider Migration

9 min read · updated August 11, 2026

The finance question is always the same: the sum of what we charged the teams is less than what we paid the provider, and it was not last month. Two causes account for nearly all of it, and both are one line of code.

The symptom: teams under-billed, invoice unchanged

The chargeback job runs nightly. It reads the request log, multiplies token counts by a price, groups by team, and writes a row per team per day. It has reconciled against the invoice for a year. In the first full month after the migration, the sum of the team rows comes in well under the invoice — and, tellingly, the shortfall is roughly stable as a fraction of spend rather than tracking any one team.

A stable fraction is the signature of a systematic under-count, not of missing rows. If requests were being dropped you would see it concentrated in whichever service failed to emit them. A flat percentage across every team means every request is being priced slightly wrong, in the same direction.

Doing the subtraction

Take one request and price it by hand. This is the fastest route to the cause and it takes about ten minutes.

Suppose the request carries a large cached system prompt. On the old provider the usage object reported a single prompt-token figure that included the cached portion, and the chargeback job multiplied it by the input rate. On the new provider, the headline input count excludes the cached tokens entirely: Anthropic’s documentation states that input_tokens represents only the tokens after the last cache breakpoint, with cache_read_input_tokens and cache_creation_input_tokens reported as separate, non-overlapping buckets. Anthropic’s prompt caching documentation gives the reconstruction as the sum of all three.

A job that reads only the headline field therefore charges the team for the tail of the prompt and nothing else. Work an example with figures you substitute from your own logs. Call the cached prefix C tokens, the per-request tail T tokens, the standard input price P per token. Assume, as a placeholder, a cache-read multiplier of one tenth of the input rate and a five-minute cache-write multiplier of one and a quarter times it — the multipliers Anthropic documents at the time of writing, which you should confirm against current pricing before relying on them.

True input cost of a cache-hit request:
    (C x 0.1P) + (T x P)

What the unchanged job charges:
    T x P

Missing per request:
    C x 0.1P

As a fraction of true input cost:
    (C x 0.1) / ((C x 0.1) + T)

With C = 20,000 and T = 500  ->  2000 / 2500 = 80% of input cost missing
With C = 4,000  and T = 500  ->   400 /  900 = 44% of input cost missing

Two things fall out of that arithmetic. The shortfall is not small, which is why it was noticed. And it is largest for exactly the teams who did the most work to make their prompts cacheable, which makes the under-charge both wrong and backwards as an incentive. Cache writes are worse still: they are billed at a premium to the standard rate, and a job reading only the headline field misses them completely.

The multipliers above are placeholders for the arithmetic, taken from the vendor’s pricing documentation at the time of writing. Substitute your own contracted rates; the shape of the result does not depend on the exact values, only on cache reads being much cheaper than fresh input and cache writes being dearer.

The second failure: the price table’s default branch

The other half of the gap is usually the price lookup. Chargeback jobs almost always contain a dictionary from model string to rate, and almost always contain a fallback for unknown models — because someone once got a KeyError in production at month end and made it stop.

That fallback is a landmine under a migration. The new model string is not in the table, so every request falls to the default, and the default is whatever seemed sensible in a hurry: the old model’s rate, a house average, or zero. If it is zero, one team’s entire spend disappears and nobody sees an error. If it is the old rate, the numbers look plausible and are wrong by the ratio between two prices, which is the hardest version to detect.

The fix is not a better default. It is no default: an unknown model string must fail the job loudly, because a chargeback number produced by guessing is worse than a chargeback number that is late. Keep a separate, explicit dated price table with effective-from rows so a mid-month price change does not retroactively reprice history. Migrating the price lookup table works through that structure.

The fix, as a diff

- input_tokens  = usage["input_tokens"]
- cost = input_tokens * PRICE.get(model, DEFAULT_RATE) + \
-        usage["output_tokens"] * PRICE_OUT.get(model, DEFAULT_RATE)

+ rates = PRICE[model]            # raises on an unknown model, deliberately
+ fresh  = usage["input_tokens"]
+ read   = usage.get("cache_read_input_tokens", 0)
+ write  = usage.get("cache_creation_input_tokens", 0)
+ out    = usage["output_tokens"]
+
+ cost = (fresh * rates.input
+       + read  * rates.cache_read
+       + write * rates.cache_write
+       + out   * rates.output)

Note what the corrected version does not do: it does not try to reconstruct a single “input tokens” number. That number is the thing whose meaning changed, and any code that keeps it keeps the bug latent for the next migration. Price each bucket at its own rate and the question of what the headline field includes stops being asked. The normalized usage schema is the version of this you build once and reuse; this page is the version you apply tonight to the job that is already wrong.

There is a third failure on the output side that surfaces a month later if you do not look for it now. Reasoning and thinking tokens are billed as output and are already inside the headline output count, so the arithmetic is right — but they are the component most likely to change size on a migration, and a chargeback report that cannot break them out cannot answer the question a team lead will ask when their line item doubles. The counts are available: OpenAI reports completion_tokens_details.reasoning_tokens, Anthropic reports output_tokens_details.thinking_tokens, and Gemini reports thoughtsTokenCount. Store them per request even though they do not change the total, because “your bill went up because the new model thinks before it answers” is a supportable statement and “output went up” is not.

What to do about the months already charged

Do not silently reprice. A chargeback figure a team has already budgeted against is a number they made decisions with, and quietly changing it costs more trust than the money is worth.

  1. Recompute the affected period from the raw request log with the corrected job, and keep both figures side by side. If you did not store the full usage object, store it from now on; this is the reason to.
  2. Reconcile the corrected total against the provider invoice for those months. Agreement to within rounding is what licenses you to talk about it.
  3. Publish the delta per team with the mechanism, before proposing any adjustment. “Cached input was not being charged” is a sentence people accept; a changed number is not.
  4. Add the invoice reconciliation as a scheduled check rather than a month-end scramble, so the next divergence is caught in days. It is the same class of check as a token count that does not match the bill.