Mapping Multi-Turn Tool-Call Sequences Between APIs
11 min read · updated August 11, 2026
A conversation containing a tool call is not one message with a field on it. It is two or three turns whose roles, nesting and correlation keys differ between APIs — and a replay that gets it wrong does not error, it silently arrives with the tool result missing.
The exchange
One exchange, written out below in three shapes. The user asks a question, the model calls a tool, your code runs it, the model answers. Four logical steps.
- User: “What is the weather in Oslo?”
- Assistant: invoke
get_weatherwith a city argument. - Your code: run it, get back a temperature and a condition.
- Assistant: “It is 4°C and overcast in Oslo.”
Steps 2 and 3 are the ones that move. Everything below is about where they go and how they are tied to each other.
The tool-message shape
OpenAI’s Chat Completions API adds a dedicated role. The assistant turn carries a tool_calls array and, on a pure tool turn, a null content; each result comes back as its own message with role tool, correlated by tool_call_id (OpenAI, Chat Completions reference).
[
{ "role": "user", "content": "What is the weather in Oslo?" },
{ "role": "assistant",
"content": null,
"tool_calls": [
{ "id": "call_abc123",
"type": "function",
"function": { "name": "get_weather",
"arguments": "{\"city\":\"Oslo\"}" } }
] },
{ "role": "tool",
"tool_call_id": "call_abc123",
"content": "{\"temp_c\":4,\"condition\":\"overcast\"}" },
{ "role": "assistant", "content": "It is 4°C and overcast in Oslo." }
]Two details catch people. The arguments are a JSON string, not a JSON object, so they must be parsed before use and serialised before replay — and a model can produce a string that is not valid JSON, which is a case your dispatcher must handle rather than assume away. And a tool message must follow an assistant message that actually contains a matching call id; a history where the assistant turn was summarised, or where the tool result was kept and the call dropped, is rejected.
The content-block shape
Anthropic’s Messages API has no tool role. Everything is a content block inside a user or assistant turn, and this is the single most important difference on this page: the tool result is carried by a message with role user (Anthropic, Messages API).
[
{ "role": "user", "content": "What is the weather in Oslo?" },
{ "role": "assistant",
"content": [
{ "type": "tool_use",
"id": "toolu_abc123",
"name": "get_weather",
"input": { "city": "Oslo" } }
] },
{ "role": "user",
"content": [
{ "type": "tool_result",
"tool_use_id": "toolu_abc123",
"content": "{\"temp_c\":4,\"condition\":\"overcast\"}" }
] },
{ "role": "assistant",
"content": [ { "type": "text", "text": "It is 4°C and overcast in Oslo." } ] }
]The differences that matter: input is a parsed object rather than a string, which is more convenient and means a converter must serialise in one direction and parse in the other. Correlation is by tool_use_id against the block’s id. The result block also accepts an is_error flag, giving a first-class way to say the tool failed — something the tool-message shape has no field for, so converting in that direction means folding the error into the content string and losing the distinction. And an assistant turn can mix a text block and a tool-use block, so a converter that assumes a tool turn has no prose will drop the prose.
The function-part shape
Google’s Gemini API uses contents rather than messages, roles of user and model rather than user and assistant, and a parts array per turn holding a functionCall or a functionResponse (Google, generateContent reference).
{ "contents": [
{ "role": "user",
"parts": [ { "text": "What is the weather in Oslo?" } ] },
{ "role": "model",
"parts": [ { "functionCall": { "name": "get_weather",
"args": { "city": "Oslo" } } } ] },
{ "role": "user",
"parts": [ { "functionResponse": {
"name": "get_weather",
"response": { "temp_c": 4, "condition": "overcast" } } } ] },
{ "role": "model",
"parts": [ { "text": "It is 4°C and overcast in Oslo." } ] }
] }The structural point worth carrying away: there is no call id. The response is correlated to the call by function name. For a single call that is fine. For two calls to the same function in one turn — the weather in Oslo and in Bergen — a name is not a unique key, and any converter that assumes an id exists has to synthesise one and remember that the wire format cannot carry it. This is the same constraint that makes a serialised parallel-call shim need its own identifiers.
What a converter loses
Now the useful part: the specific places a history converter drops something, ranked by how quietly it does it.
- The role of the result turn. A converter written against the tool-message shape looks for role
tool. Against a content-block history there is none — the results are insideuserturns — so a filter for tool messages returns nothing and the results vanish. The model then sees itself calling a tool and immediately answering, with no data. It will usually answer anyway, confidently, which is why this is the worst failure here. - Arguments as string versus object. Serialise an already-serialised argument string and you get a JSON-encoded string where an object was expected. It parses; it is wrong.
- The error flag. There is no counterpart for
is_errorin the tool-message shape. Encode it in the content in a form the model can read — a short string saying the tool failed and why — and accept that the machine-readable distinction is gone. - Text alongside a tool call. An assistant turn that reasons in prose and then calls a tool has two blocks. A converter that takes only the first, or only the tool block, changes the visible conversation.
- Non-text tool results. Some APIs allow a tool result to contain an image block; others accept a string only. Converting towards the stricter side, the image cannot be represented in the result and must be attached elsewhere or dropped.
- Unanswered calls. Every invocation in an assistant turn generally must be answered before the conversation continues. Filtering a history — dropping old turns to fit a context window is the usual reason — can leave a call whose result was trimmed, or a result whose call was trimmed. Both are rejected, and the error names a message index in an array you built programmatically, which is a miserable thing to debug. Trim in whole call-and-result units.
Store one canonical form
The structural fix is to stop treating any provider’s wire format as your storage format. If your database holds OpenAI-shaped messages because that is where you started, every other provider is a lossy conversion from a format that cannot express is_error, and every new provider is another pairwise converter.
Define your own turn type — role, an ordered list of parts, each part text or invocation or result, invocations and results carrying your own generated id, results carrying an explicit failure flag — and write one serialiser per provider from it. The canonical form should be a superset of what every provider expresses, so that conversion loses things only at the wire and never in your own store. Then a conversation recorded against one provider can be replayed against another, which is what makes testing multi-turn tool calling across providers possible at all.
Two rules keep it honest. Generate your own call ids rather than storing the provider’s, so a history is not tied to the API that produced it. And store the raw provider response alongside the canonical form for a retention window, because the first time a conversion is wrong you will need to see what actually came back, and a canonical record cannot show you what it failed to capture.