Skip to content

GPU Cost per Million Tokens

Rental or purchase, power and facility overhead divided by the tokens a GPU actually produces, with the cost per 1M shown across six utilisation levels.

What the hardware produces

If you rent

If you buy

Rented, at 40.0% utilisation
$1.93 per 1M output tokens

Owned outright, the same hardware at the same utilisation comes to $1.24 per 1M — 0.64× the rental figure. At 100% utilisation the rented figure would be $0.77.

Output tokens per month at this utilisation
946M
Rental cost per month
$1,825
Depreciation per month
$833.33
Electricity per month
$85.85 (715 kWh)
Rack and network per month
$250.00
Owned cost per month, all in
$1,169
Rented, per 1M output tokens
$1.93 / 1M
Owned, per 1M output tokens
$1.24 / 1M
Rented, at 10% utilisation
$7.72 / 1M
Rented, at 25% utilisation
$3.09 / 1M
Rented, at 40% utilisation
$1.93 / 1M
Rented, at 60% utilisation
$1.29 / 1M
Rented, at 80% utilisation
$0.96 / 1M
Rented, at 100% utilisation
$0.77 / 1M
What this assumes: throughput is output tokens per second — the decode phase, which is memory-bandwidth-bound and is what limits a serving box. Prefill runs far faster per token and is excluded, so a workload with very long prompts and short answers will do better than this suggests. The owned figure is straight-line depreciation with no residual value, plus measured-style power at the wattage and PUE you supply, plus a flat facility charge; it excludes the capital cost of the money, the host machine around the GPUs, spares, and the people who rack it. Utilisation is applied to powered-on hours: a GPU that is up but idle costs exactly as much as one that is saturated. Every figure here is a field you typed, including the defaults.

The sensitivity rows are the point of this page. Everything else — hardware price, electricity, whether you rent or buy — moves the answer by tens of per cent. Utilisation moves it by multiples. A GPU at 10% utilisation costs ten times as much per million tokens as the same GPU at 100%, and no procurement decision available to you has that kind of leverage. Comparisons that quote a cost per million from peak throughput are quoting the 100% row and calling it the answer.

The reason real utilisation is low is structural rather than a failure of engineering. You size for peak traffic, traffic is diurnal, and the hardware is billed by the hour regardless. Continuous batching helps a great deal because it keeps the batch full as requests arrive and finish at different times. Consolidating several models or several teams onto one fleet helps more. Filling the overnight trough with asynchronous work — evals, re-embedding, backfills — helps most, because it converts idle hours into output at no additional hardware cost at all.

Rent versus buy is a smaller question than it looks. Buying wins on steady-state cost per hour and loses on flexibility, and the depreciation period you choose largely determines the answer: the same card over 24 months versus 48 is a two-fold difference in the monthly figure, and the honest life of an accelerator depends on when something meaningfully faster ships rather than on when it breaks. Whichever way you go, the number that will actually determine your cost per token is the utilisation field, and it is the one nobody measures before signing.

GPU Cost per Million Tokens · Multigrid