Claude's Citations Feature: How Source Spans Come Back in the Response
8 min read · updated August 11, 2026
Citations turn “the model says the contract permits it” into a character range in a specific document. The feature is one flag on a document block; the interesting part is what it does to the shape of the response.
Enabling it on a document
Citations attach to document content blocks, not to the request as a whole. Set citations on each document you want cited:
{
"model": "claude-opus-4-6",
"max_tokens": 2048,
"messages": [
{
"role": "user",
"content": [
{
"type": "document",
"title": "Master Services Agreement",
"source": {
"type": "text",
"media_type": "text/plain",
"data": "1. Term. This Agreement begins on the Effective Date…"
},
"citations": {"enabled": true}
},
{
"type": "text",
"text": "Under what conditions can the customer terminate early?"
}
]
}
]
}The source may be plain text as above, a base64 PDF, a file reference, or an array of content blocks. The document title is worth setting even though it is optional, because it comes back on every citation and is what you will render.
The flag is all-or-nothing across the documents in a request: enable it on some and not others and the request is rejected. Decide once per request.
What the response looks like
This is the part that breaks existing code. Without citations, an answer is usually one text block. With citations, the response is split into several text blocks — cited spans and uncited connective prose — and only the cited ones carry a citations array:
{
"content": [
{
"type": "text",
"text": "The customer may terminate early "
},
{
"type": "text",
"text": "for material breach that remains uncured after thirty days' notice",
"citations": [
{
"type": "char_location",
"cited_text": "Customer may terminate this Agreement upon a material breach by Provider that remains uncured thirty (30) days after written notice.",
"document_index": 0,
"document_title": "Master Services Agreement",
"start_char_index": 4187,
"end_char_index": 4321
}
]
},
{
"type": "text",
"text": ", and on ninety days' notice for convenience."
}
],
"stop_reason": "end_turn"
}So content[0].text is no longer the answer — it is the first fragment of it. Rendering means walking the whole array, concatenating the text, and decorating the fragments that have citations attached. Code that reads the first block and stops will display a partial sentence.
Every citation object carries cited_text, which is the source text verbatim rather than the model’s paraphrase of it. That is the field that makes the feature worth using: you can display the quotation without a second retrieval step, and you can verify the model’s claim against it programmatically.
The three location types
The type field on a citation names how the span is addressed, and it depends on what kind of document was supplied:
char_location— for plain text documents. Carriesstart_char_indexandend_char_indexinto the text you supplied. Exact enough to highlight in place.page_location— for PDFs. Carriesstart_page_numberandend_page_number, and the numbering is 1-indexed. Off-by-one errors here are the most common bug in citation rendering, because almost every PDF library is 0-indexed.content_block_location— for documents supplied as an array of content blocks. Carriesstart_block_indexandend_block_index, addressing your own chunks rather than characters.
All three also carry document_index — the position of the document in the content array — and document_title. With several documents attached, document_index is what tells you which one the span came from.
On a streamed request the citation does not arrive with the text it annotates. The text streams as ordinary text_delta events, and the citation arrives separately as a citations_delta on the same block index. A renderer that decorates text as it arrives therefore has to be able to attach a citation to a span it has already painted — or, more simply, hold each block until its content_block_stop before rendering it. Streaming a citation-enabled answer word by word and decorating it correctly at the same time is genuinely harder than the buffered case, and it is worth deciding which of the two you want before building either.
Why prompting for citations does not work
The workaround everyone tries first is to skip the feature and ask for it in words: paste the document into the prompt, and instruct the model to quote the source and give the character offsets. It produces output that looks right, and it is unreliable in a specific and instructive way.
The quotations are usually close and not exact. A model reproducing a passage from memory of its own context will normalise a hyphen, drop a parenthetical, or tidy the grammar — so a verification step that searches your source for the quoted string fails on text that was substantively correct, and you cannot distinguish that from a genuine fabrication.
The offsets are worse. A character index is a counting task over thousands of characters, and it is exactly the kind of work a language model is poorest at. Numbers come back confidently and are wrong, and because they are plausible, a highlight built on them lands near the right passage often enough that the bug survives review.
The feature is not doing the same thing more conveniently. The spans are produced by the serving infrastructure against the document as supplied, so cited_text is the source text rather than a reproduction of it, and the indices are computed rather than estimated. That is the reason to use it: not the formatting, but the fact that the answer to “does this quotation actually appear in the document” stops depending on the model.
Controlling granularity with custom content
Supplying a document as an array of content blocks is the way to control what a citation can point at. Each block becomes an independently citable unit, so if your retrieval pipeline already has chunks with identifiers, you can make citations line up with them exactly:
{
"type": "document",
"title": "Support knowledge base",
"source": {
"type": "content",
"content": [
{"type": "text", "text": "KB-1041: Password resets expire after 60 minutes."},
{"type": "text", "text": "KB-1042: SSO users cannot reset passwords locally."}
]
},
"citations": {"enabled": true}
}A citation then comes back as a content_block_location with a block index you can map straight back to KB-1041 — no character-offset arithmetic, no fuzzy matching of cited_text against your corpus.
Chunk size is the tradeoff. Large blocks give citations that are easy to map and imprecise to display — a citation that points at a 2,000-word section is barely more useful than naming the document. Small blocks give precise highlights and more of them, and they cost the model more effort to assemble an answer from. Paragraph-sized blocks are the usual compromise, and if you already chunk for embedding retrieval, reusing those boundaries means a citation maps directly onto a row in your index.
Constraints worth knowing first
- Incompatible with structured output. Combining citations with a forced JSON response format is rejected. The two features both control the shape of the output and the conflict is not resolvable, so if you need cited claims inside a JSON envelope you assemble it yourself from the block array.
- The model decides when to cite. Enabling the feature makes citation possible, not mandatory. An answer drawing on general knowledge rather than the document will have no citations attached, which is correct behaviour and something your UI has to represent.
- Documents are still tokens. Citations change how the answer is annotated, not how the input is priced. A 200-page PDF costs what a 200-page PDF costs on every turn it stays in the conversation.
The reference is Anthropic’s citations documentation, which lists the supported source types and the current field set.