Claude and Multiple Images in One Request: Ordering Rules
8 min read · updated August 11, 2026
Claude accepts many images in a single request, and the number is higher than most people assume. What is not obvious is that images have no names — position in the content array is the only thing distinguishing one from another, which changes how you have to write the prompt around them.
The documented limits
From Anthropic’s vision documentation, at the time of writing:
- Up to 100 images per request through the API. The claude.ai apps allow far fewer — 20 per turn — so a prompt that works in your code will not necessarily paste into the chat interface.
- 5 MB per image when supplying base64 through the API.
- 8,000 × 8,000 pixels maximum per image, and when a request contains more than 20 images the per-image ceiling drops to 2,000 × 2,000 pixels.
- Images are resized down automatically if they exceed what the model uses, but they are never scaled up. Sending a 4K screenshot buys nothing over sending a correctly downscaled one, and costs upload bandwidth.
The context window is the other limit, and often the binding one. A hundred images do not fail on the image count — they fail on tokens. Anthropic documents an approximation of tokens ≈ (width × height) / 750 for image input, which puts a single 1,000 × 1,000 image at roughly 1,300 tokens. Fifty of those is around 65,000 tokens before a word of your prompt. The arithmetic and the exact rules are on Claude’s image token cost.
Worth separating from the count limit: an image occupies context for the rest of the conversation, and the count limit is per request rather than per conversation. Ten images sent across five turns are ten images in context on the fifth turn, all of them re-sent and all of them charged. The per-request ceiling will not stop you assembling a conversation that exceeds the window; only counting will.
Position is the only identifier
A message’s content is an ordered array, and image blocks sit in it alongside text blocks. There is no id, name or alt field on an image block. The model sees them in the order you wrote them, interleaved with whatever text surrounds them, and that interleaving is the entire addressing mechanism:
{
"role": "user",
"content": [
{ "type": "text", "text": "Image 1 — the design:" },
{ "type": "image",
"source": { "type": "base64", "media_type": "image/png", "data": "iVBORw0..." } },
{ "type": "text", "text": "Image 2 — the current build:" },
{ "type": "image",
"source": { "type": "url", "url": "https://example.com/build.png" } },
{ "type": "text",
"text": "List every visual difference between image 1 and image 2, grouped by component." }
]
}Two consequences follow. First, if you place both images together and then ask your question, the model has to infer which is which from content alone, and “the first image” becomes a guess it can get backwards. Second, if your application builds the array from a set or a map rather than a list, the ordering can vary between requests while the prompt text still says “image 1” — a bug that reproduces intermittently and looks like model unreliability.
Labelling images so text can refer to them
The reliable pattern is to precede each image with a short text block naming it, and then use those names in the instruction:
- Emit a text block per image containing nothing but the label —
"Image 1: invoice, page 1". Keep the labels short; they are prompt tokens. - Emit the image block immediately after its label.
- Put the instruction in a final text block after all the images, referring to them by the labels you used. An instruction before the images is answered with less of the evidence in view.
- If the answer must be machine-readable per image, ask for the label as a field:
"Return one JSON object per image with keys image, total, date". Now the join back to your own files does not depend on the model preserving order in its reply.
Sources can be mixed freely: some blocks base64, some URL-referenced, some referring to previously uploaded files. The addressing rule is the same regardless of source, and the media type must be one of the supported ones — JPEG, PNG, GIF and WebP — declared accurately in media_type. A PNG declared as image/jpeg is rejected rather than sniffed.
Images can also be split across several user messages rather than crammed into one, which is the better structure for a comparison built up over a conversation. The addressing rule survives that: what the model has is a single ordered sequence, and a label in the message that introduced an image still refers to it three turns later. What does not survive is renumbering — if turn four introduces a new “image 1”, you now have two, and the model will pick one. Number across the conversation, not within the message.
How multi-image requests fail
Four failure modes, in roughly the order you will meet them.
- A declared media type that does not match the bytes.
media_typeis taken at its word for validation and then contradicted by the decoder, so a PNG sent asimage/jpegis rejected rather than sniffed and corrected. This is the most common bug in code that derives the media type from a filename extension, because a great many files ending in.jpgare not JPEGs. - Base64 with a data-URL prefix still attached. The
datafield takes the encoded bytes alone. Leavingdata:image/png;base64,on the front produces a decode error, and it happens whenever a payload is copied straight out of a browserFileReaderresult. - A URL the API cannot fetch. URL-sourced images are fetched server-side, so anything behind authentication, behind a firewall, on localhost, or gated by a hotlink referer check will fail — even though it loads perfectly in your browser. If the asset is private, upload it or send the bytes.
- Too many tokens rather than too many images. Past a certain point the failure is a context-length error naming a token count, which sends people looking for an image-count limit that is not what they hit.
There is a fifth, quieter failure that is not an error: the model answering about the wrong image. Almost always the cause is either an unlabelled batch of images, or a client that built the content array from an unordered collection. If a multi-image prompt is right most of the time and wrong occasionally, check that the array order is deterministic before you change a word of the prompt.
What multiple images cost you
Three costs that are easy to miss when a multi-image prompt is being designed:
- Tokens scale with area, not with count. Ten small images can be cheaper than one large one. Downscaling to the smallest size at which the detail you are asking about is still legible is the single biggest lever, and it is applied before upload, not by a parameter.
- Every image is re-sent on every turn. There is no server-side conversation state, so a ten-image comparison that runs to five turns pays for those images five times. This is the case prompt caching is built for: a cache breakpoint after the image blocks turns the repeat cost into the cached-read rate.
- Failure is at the window, not the count. Too many images produces a context-length error rather than an image-count error, so the message you get will talk about tokens. That is covered in the context window exceeded error.
For documents specifically, sending pages as images is usually the wrong route — PDF input is handled natively and accounts for tokens differently, which is the subject of Claude’s PDF input tokens.