Skip to content

Prefilling Claude’s Assistant Turn: How It Changes the Response

8 min read · updated August 11, 2026

Claude’s Messages API does not require the conversation to end on a user turn. If the last message in the array has role: "assistant", the model treats its content as the beginning of its own reply and continues from there — which constrains the response far more reliably than asking politely in the prompt.

The mechanism

There is no special parameter. You simply append an assistant message with the text you want the reply to start with:

{
  "model": "claude-sonnet-4-20250514",
  "max_tokens": 1024,
  "messages": [
    { "role": "user", "content": "Extract the invoice fields from this email:\n\n..." },
    { "role": "assistant", "content": "{" }
  ]
}

Why this works is worth being precise about, because it is not a feature bolted on top of the model. A chat model is still a next-token predictor over a single sequence; the roles are formatting inside that sequence. Ending the sequence part-way through the assistant’s turn means the highest-probability continuations are the ones that complete what is already there. An opening brace makes JSON the overwhelmingly likely continuation, because the training data contains almost nothing where a brace at the start of a reply is followed by “Sure! Here’s the data you asked for:”.

That is the difference between prefilling and instructing. An instruction competes with every other pressure on the distribution. A prefill changes what the model is completing.

The prefill is not echoed back

This is the detail that breaks integrations, and it is easy to miss when you are reading the model’s output in a terminal. The response contains only the continuation. Your prefilled text is not repeated in the returned content.

// You sent, as the final message:  { "role": "assistant", "content": "{" }
// You get back:

{
  "content": [
    { "type": "text",
      "text": "\"invoice_number\": \"INV-2041\", \"total\": 412.50 }" }
  ],
  "stop_reason": "end_turn"
}

// The complete JSON document is prefill + response text:
//   "{" + "\"invoice_number\": ... }"

So JSON.parse(response.content[0].text) fails, and it fails with a message about an unexpected token that points nowhere useful. The correct reconstruction is to concatenate your prefill onto the front of the returned text before parsing. Keep the prefill string in a variable and use the same variable in both places rather than typing the brace twice.

The same asymmetry applies to usage: the prefilled tokens are billed as input, not output, because they were part of the request.

What it is actually good for

  • Suppressing preamble. Prefilling a single character of the expected output removes the introductory sentence far more reliably than “respond with only the JSON” in the system prompt.
  • Forcing a format. An opening brace forces JSON, a triple backtick and language tag forces a code block, an opening XML tag forces the matching structure. Pairing this with stop_sequences — for example stopping on "```" after prefilling one — gives you exactly the fenced body and nothing around it.
  • Continuing a truncated response. When stop_reason came back as max_tokens, appending the partial text as an assistant message and calling again resumes from the cut point instead of restarting.
  • Holding a persona or a schema key order. Prefilling "{\"reasoning\":" both forces JSON and forces the reasoning field to come first, which matters because a field generated before the answer can influence it and a field generated after it cannot.

For strict schema conformance rather than merely “valid JSON”, a forced tool call is the stronger mechanism — the schema is validated rather than hoped for. That trade-off is covered in getting JSON out of Claude with tool_choice. Prefill is the lighter option when you want shape without wiring up a tool.

Two patterns worth copying

The fenced-block extractor. Prefill the opening fence and stop on the closing one. The response is then the body and nothing else — no preamble, no trailing explanation, no fence to strip:

{
  "model": "claude-sonnet-4-20250514",
  "max_tokens": 2048,
  "stop_sequences": ["```"],
  "messages": [
    { "role": "user", "content": "Rewrite this function to use async/await:\n\n..." },
    { "role": "assistant", "content": "```typescript\n" }
  ]
}

Two things to notice. The prefill ends with a newline inside the string, which is allowed — the rule is about the content ending in whitespace, so a fence followed by a newline is a common trip hazard; if the API rejects it, end the prefill at the language tag instead and let the model emit the newline. And stop_reason will come back as stop_sequence rather than end_turn, which is the signal that the block closed properly rather than the model running out of room.

The resume. When a response ends on max_tokens, the partial text becomes the prefill of the next request:

let text = "";
let stop = null;

do {
  const messages = [{ role: "user", content: prompt }];
  if (text) messages.push({ role: "assistant", content: text.trimEnd() });

  const res = await client.messages.create({
    model: "claude-sonnet-4-20250514",
    max_tokens: 4096,
    messages,
  });

  text += res.content.filter(b => b.type === "text").map(b => b.text).join("");
  stop = res.stop_reason;
} while (stop === "max_tokens");

The trimEnd() is not cosmetic — it is what keeps the trailing whitespace rule from failing the second request, and it is the line this loop is always missing the first time somebody writes it. Note also that each iteration re-sends the growing text as input, so a document assembled over five resumes pays for the prefix five times. Below a certain length it is cheaper to raise max_tokens; above it, chunk the work by section instead of resuming.

Where it does not work

Four documented restrictions, each of which produces a different symptom:

  • No trailing whitespace. The final assistant content may not end with a space or newline. The API returns HTTP 400 with a message along the lines of final assistant content cannot end with trailing whitespace. Prefill "{", not "{ ". This is the single most common prefill error and it is entirely mechanical: the tokeniser would otherwise have to decide between a whitespace-prefixed token and a bare one, and the API refuses rather than guessing.
  • Not with extended thinking. A thinking-enabled request cannot end on an assistant message, because the model must produce its thinking block first and the prefill would precede it. Requests that try are rejected.
  • Not on every surface. Prefill is a Messages API behaviour. It is not available through interfaces that always append a fresh assistant turn for you, and third-party wrappers that normalise to an OpenAI-shaped chat array frequently drop or reorder a trailing assistant message.
  • It biases more than the first token. A long prefill is a strong constraint on register and content, not just format. Prefilling several sentences of an argument and asking for balance is asking for two contradictory things.

One further caution that only shows up over time: prefill behaviour is a property of the model, not of the API surface, and it is not versioned separately from the model. A prefill that reliably produced bare JSON on one model can produce JSON wrapped in a code fence on its successor, because the successor was trained with different habits around formatting. Nothing errors; the parse just starts failing. If you depend on prefill for machine-readable output, the model upgrade checklist needs a case for it, and the check is the parse rather than a human reading the output — a fence is invisible when you are eyeballing a terminal and fatal when you are calling JSON.parse.

The same applies to prefill through anything that is not the Messages API directly. Wrappers, framework abstractions and gateways all have to decide what to do with a trailing assistant message, and the safe ones pass it through untouched. If a prefill that works against the raw API stops working through a layer, the first thing to check is whether that layer appended its own turn or merged your assistant message into the preceding user one.

Prefill is also, unavoidably, a way to push a model past a refusal, and Anthropic’s usage policies apply to it exactly as they apply to the prompt. A refusal that has been prefilled around is not a signal that the request became acceptable.