Skip to content

Cost Report Generator

Turn a usage export and your own rates into a shareable markdown or CSV cost report, broken down by whatever the rows are named after.

Total for July 2026
$339.80

4 line(s) priced, 326,370 requests. Rates are the ones you typed; nothing here is read from a price list, so this is your arithmetic rather than anyone else's.

Cost per request, averaged
$0.0010
Largest line
chat-assistant — $182.73
Input (uncached + cached)
$101.62
Output
$238.19
Output share of the bill
70.1%
What this assumes: rates are per 1,000,000 tokens and are yours — typed into the box, never looked up, never shipped with this page. Cached input tokens are treated as a SUBSET of the input tokens rather than an addition, which is how the major providers report them; if yours reports them separately, subtract them from the input column before pasting or the total will be too high. Every usage row whose name has no rate line is excluded from the total and named in the list above, because a cost report that quietly prices something at zero is worse than one that admits a gap. Token costs only: no images, no audio, no storage, no fine-tuning, no per-hour anything.

The monthly question is never “what did we spend” — the provider's dashboard answers that. It is “what did we spend it on, and which line moved”, and that needs your usage broken down by the thing you actually care about: a feature, a customer, a job. This turns that export into something you can paste into a document and defend.

The two numbers that change the conversation

Cost per request and the output share. Cost per request is the one that scales: a feature at four thousandths of a dollar a call is fine at ten thousand calls and a budget item at ten million. The output share tells you which lever to pull — a bill that is 70% output tokens is a bill about how much the model writes, and no amount of prompt trimming will touch it, while a bill that is 90% input is a bill about context you are resending and caching may cut it in half.

Cached tokens are the usual arithmetic error

Providers report cached input tokens as part of the input total, not on top of it. Adding them produces a report that overstates the bill by however much your caching is working, which is exactly backwards. This page subtracts them and prices the remainder at the full rate and the cached portion at the cached rate — and if you gave no cached rate, it says so rather than assuming a discount you may not have.

What a report like this leaves out

Token costs are rarely the whole bill. Embeddings you regenerate, vector storage, the retries that failed and were still charged, image and audio units billed by the item, fine-tuning runs, per-hour endpoints — none of that is in a token export, so none of it is here. Say so in the document you paste this into. A cost report that is understood to be partial is useful; one that is assumed to be complete is how a surprise gets planned for six months in advance.

Cost Report Generator · Multigrid