PDF Input in the Claude API: How Pages Become Tokens
8 min read · updated August 11, 2026
A PDF sent to the Claude API is not read as a file. Each page is turned into two things the model already understands, and the token cost is the sum of both — which is why the documented per-page figure is a range rather than a number.
What happens to a page
Anthropic documents the handling explicitly: for each page of a submitted PDF, the text is extracted, and the page is also rendered as an image. Both are placed in the context. The model therefore sees the words twice over — once as tokens it can quote exactly, and once as pixels that carry the layout, the table rules, the signature, the chart nobody wrote out in prose.
That is a design decision with a clear motive. Text extraction alone loses the two-column reading order, the fact that a number sat in the “2025” column, and anything drawn rather than typed. Image alone loses exactness: a model reading pixels can transcribe an account number wrong. Sending both costs more and fails less.
It also means PDF cost is not a separate pricing line. There is no per-page fee. You are billed input tokens, at the model’s ordinary input rate, for the extracted text and for the rendered page image.
Why the range is so wide
Anthropic’s documentation estimates roughly 1,500 to 3,000 tokens per page. A factor of two is unusually vague for a documented figure, and the reason is visible once you split the number in half.
The image half is close to fixed. A rendered page is downscaled to the same ceiling as any other image — around 1,600 tokens at the top, typically somewhat under it for a page rendered at a sensible resolution. Whatever is on the page, that contribution barely moves. The arithmetic behind that ceiling is on the image token cost page.
The text half is where all the variance lives. A title page with a logo and eight words contributes almost nothing. A dense page of a legal agreement in nine-point type might carry 700 words, which is on the order of 900 to 1,000 tokens of English. Add the two:
title page ≈ 1,500 (image) + ~20 (text) ≈ 1,520 tokens prose page ≈ 1,500 (image) + ~450 (text) ≈ 1,950 tokens dense contract ≈ 1,500 (image) + ~1,000 (text) ≈ 2,500 tokens dense + tables ≈ 1,500 (image) + ~1,400 (text) ≈ 2,900 tokens Assumption: English prose at roughly 0.75 words per token, and a rendered page landing near the documented per-image ceiling. Both are estimates; count_tokens is the authority for a given file.
The documented 1,500–3,000 range is exactly the span between an almost empty page and a full one, and it tells you which end of it your own documents sit at. Scanned pages with no extractable text layer sit near the bottom of the range in tokens, which is not good news: they cost less because the model is getting less, and the extraction it would have relied on for exact figures is absent.
That last point is worth stating as a rule, because it inverts the usual relationship between price and quality. On a born-digital PDF — one exported from a word processor or a reporting tool — the text layer is exact, and a figure the model quotes back to you came from characters rather than from pixels. On a scan or a photographed document, there may be no text layer at all, and every number the model reports is an act of visual transcription at roughly 1.15 megapixels. The token cost barely tells the two apart. If your source documents are scans and the answers have to be exact, the mitigation is an OCR pass before the model sees the file, so that a text layer exists to be extracted, not a larger model.
It is also why the range should not be used as a per-page constant in a cost model without checking which end you are at. A pipeline budgeted at 1,500 tokens a page that turns out to process dense contracts will run at double its forecast, and the difference is not a rounding error across a hundred thousand pages.
A twenty-page report, priced
Take a twenty-page quarterly report: a cover, two pages of contents and summary, fourteen pages of prose and charts, three pages of dense tables.
1 cover × 1,520 = 1,520
2 summary pages × 1,900 = 3,800
14 prose + charts × 2,100 = 29,400
3 table pages × 2,900 = 8,700
---------
total ≈ 43,420 input tokens
At $3.00 per 1,000,000 input tokens (Sonnet-class):
43,420 × $3 / 1,000,000 = $0.13 per pass
Plus your prompt. A 600-token instruction is noise here; a 20,000
token few-shot preamble is a third of the bill again.Thirteen cents to read a twenty-page report once. The number that matters is not that one, it is the one after it: asking five questions about the same report in five separate requests costs sixty-five cents, because the whole document is re-sent every time. There is no session, no uploaded-file handle that the model remembers. Every request re-reads everything.
This is the case prompt caching exists for. A document sent as a cached prefix is written to the cache once and read back at a reduced rate on subsequent requests within the cache lifetime, which turns five questions from five full reads into one read and four cheap ones. See cache_control on the Messages API for how the breakpoints are placed.
The page count where it stops fitting
Anthropic documents a per-request page limit of 100 pages and a request size limit in the tens of megabytes. The page limit is not the constraint you hit first. Work it through against a 200,000-token context window:
100 pages × 1,500 tokens/page = 150,000 tokens fits, just 100 pages × 2,000 tokens/page = 200,000 tokens at the limit 100 pages × 3,000 tokens/page = 300,000 tokens rejected
So a hundred-page document is inside the documented page limit and outside the context window if its pages are dense. The failure comes back as a request-too-long error with the real token count in it, not as a page-limit error, which is why the error surprises people who checked the page limit and thought they were safe. The shape of that response is on the context window exceeded page.
The safe working assumption is roughly 60 to 90 dense pages per request, not 100, and less than that once you account for your own prompt and for the output you are reserving with max_tokens — the window has to hold both.
Making it cheaper
- Send fewer pages. Split the PDF and send the pages that could contain the answer. A page-selection step using a cheap model, or plain text search over an extracted text layer, removes most of the document before the expensive model sees it.
- Send text only, where layout does not matter. If you extract the text yourself and send it as an ordinary text block, you pay the text half and not the image half — roughly a quarter of the cost for a dense page, and much less for a sparse one. You lose layout, charts and anything drawn. That trade is right for contracts you have already parsed and wrong for invoices.
- Cache the document, not the question. Put the PDF at the start of the prompt and the varying question at the end, so the cached prefix is as long as possible.
- Count before you send. The same message array goes to
/v1/messages/count_tokensand comes back with the number. For a batch job, count one representative file and multiply. - Use the batch path for anything not interactive. Anthropic offers asynchronous message batching at a reduced rate for work that does not need an answer within seconds. Document processing is the archetypal case: nobody is watching, the deadline is measured in hours, and the discount applies to the whole of a large input bill.
The order of those matters. Sending fewer pages is a change in kind and the others are changes in degree — a retrieval step that reduces a hundred-page filing to the four pages that mention a clause takes the bill down by more than an order of magnitude, and no combination of caching and batching on the full document approaches that. Reach for the discounts after the document is as small as it can be, not instead of making it smaller.
The counter-argument, which is sometimes right, is that a retrieval step is a second thing to build, tune and debug, and its failure mode is silent: it drops the page containing the answer and the model confidently answers from what remains. Sending the whole document has no such failure. For a low-volume workload where thirteen cents a document is irrelevant, paying for the simple thing is the correct engineering decision, and it stops being correct somewhere around the point where the cost line becomes visible on an invoice.