Skip to content

Unit Testing a Tool-Calling Loop Without Calling the Model

11 min read · updated August 11, 2026

A tool-calling loop is ordinary control flow: call, branch on the response, execute, append, call again, stop. None of it needs a model, and testing it against one buys nothing except non-determinism, latency and a bill.

The seam: script the model

Define the loop against a narrow interface rather than the provider SDK — one function that takes messages and returns an assistant turn. That single indirection is what makes everything below possible, and it is worth the refactor if you do not have it.

export type AssistantTurn =
  | { kind: "text"; content: string }
  | { kind: "tool_calls"; calls: { id: string; name: string; arguments: string }[] };

export type Model = (messages: Message[]) => Promise<AssistantTurn>;

export async function runLoop(
  model: Model,
  tools: Record<string, (args: any) => Promise<unknown>>,
  messages: Message[],
  maxTurns = 8,
): Promise<Message[]> { /* ... */ }
import { expect, test, vi } from "vitest";

function scriptedModel(...turns: AssistantTurn[]) {
  const seen: Message[][] = [];
  let i = 0;
  const fn = vi.fn(async (messages: Message[]) => {
    seen.push(structuredClone(messages));
    if (i >= turns.length) throw new Error("loop called the model more times than scripted");
    return turns[i++];
  });
  return Object.assign(fn, { seen });
}

structuredClone on the way in matters: the loop mutates its message array, so storing the reference gives every recorded call the same final contents and every assertion about intermediate state becomes meaningless. Throwing when the script runs out is the other half — an infinite loop then fails with a sentence rather than hanging until the runner gives up.

The message shapes the loop must produce

This is where loops break, and it is provider-specific, so assert it exactly rather than trusting that it looks right.

OpenAI-shaped. The assistant turn is appended verbatim, including its tool_calls array. Then one message per call, with role: "tool", a tool_call_id matching the id from the request, and the result serialised into content as a string. Omitting the id, or sending one result for two calls, produces a 400 from the API rather than a wrong answer — which is good, but only if a test finds it before a user does.

Anthropic-shaped. The assistant turn contains a tool_use content block with an id, a name and an input object. The result goes back as a user message whose content is a tool_result block carrying tool_use_id. The role is the trap: it is a user message, not a tool one, and code ported from the other shape gets this wrong.

Assert the whole appended sequence in one comparison, not field by field. A single toEqual against the expected message array is more readable and fails with a usable diff:

test("a tool result is appended in the shape the API requires", async () => {
  const model = scriptedModel(
    { kind: "tool_calls", calls: [{ id: 'call_1', name: "get_weather", arguments: '{"city":"Paris"}' }] },
    { kind: "text", content: "It is 14 degrees in Paris." },
  );
  const getWeather = vi.fn(async () => ({ tempC: 14 }));

  const out = await runLoop(model, { get_weather: getWeather }, [
    { role: "user", content: "weather in Paris?" },
  ]);

  expect(getWeather).toHaveBeenCalledWith({ city: "Paris" });   // parsed, not a string
  expect(out.at(-2)).toEqual({
    role: "tool",
    tool_call_id: 'call_1',
    content: JSON.stringify({ tempC: 14 }),
  });
  expect(model).toHaveBeenCalledTimes(2);
});

Control-flow assertions

  • Termination. A text turn ends the loop. Assert model was called exactly twice for a one-tool conversation — the count is the assertion, and it catches a loop that calls once more with nothing to say.
  • The turn cap. Script a model that returns a tool call every time and assert the loop stops at maxTurns, with an error or a defined give-up result rather than by exhausting the script. Without this a mis-prompted model becomes an unbounded spend.
  • Parallel calls. A single turn can contain several calls. Assert all of them execute, that every result is appended before the next model call, and that the results are matched to the right ids — running two tools concurrently and appending in completion order rather than request order is a real and subtle bug. Assert the ids in the appended messages, not the order of execution.
  • An unknown tool name. The model can hallucinate a tool. The loop must append an error result the model can recover from, not throw. Assert the loop continues and that the error text reaches the model.
  • A tool that throws. Same principle: the exception becomes a tool result describing the failure. Assert the raised error never escapes runLoop, and that nothing sensitive from the stack trace is in the message sent back.
  • Arguments that are not valid JSON. The arguments field is a string produced by a model and it can be malformed. Assert the loop reports that as a tool error rather than crashing on JSON.parse.
  • No network happened. Run the suite with MSW’s onUnhandledRequest: "error", or assert a fetch spy was never called. This is what keeps the test a unit test as the code around it grows.

For the broader question of what belongs in a mocked loop test versus an evaluation against a real model, see testing without the model and what to do when a tool call does not fire.

The streaming variant

If the loop consumes a stream, one more layer needs its own test, because tool calls arrive in pieces. In an OpenAI-shaped stream the deltas carry a tool_calls array where each entry has an index, and the id and function.name appear once while function.arguments accumulates across many deltas. The accumulator must key on index, not on array position within a single delta, and must concatenate rather than overwrite.

Test that accumulator separately from the loop: feed it a recorded sequence of deltas and assert the reconstructed call list. Then feed the loop the reconstructed calls. Two small tests beat one large one here, because a failure tells you which half is wrong. The fragment accumulation itself is the subject of testing partial JSON mid-stream.

Writing the suite

  1. Extract the Model interface if you do not have it, and make the production code depend on it rather than on the SDK.
  2. Write scriptedModel once and put it in a test helper file. Every case below is then three or four lines.
  3. Write the happy path with one tool call, asserting the appended message shape and the call count.
  4. Add the failure cases: unknown tool, throwing tool, malformed arguments, turn-cap exhaustion. Each is a different scripted model and the same assertions.
  5. Add the parallel-call case with two calls in one turn, asserting id matching rather than execution order.
  6. If you stream, add the delta accumulator test with a recorded sequence copied from a real response rather than one you invented.