Skip to content

Claude's Computer Use Tool: How Screenshots and Actions Round-Trip

9 min read · updated August 11, 2026

Computer use is not a separate API. It is one Anthropic-defined tool whose results happen to be images, driven by exactly the same tool_use and tool_result blocks as any other tool — and the whole of it is a loop you run.

Declaring the tool

The computer tool is schema-less: you declare a type and a name and the input schema is built into the model. What you do supply is the screen geometry, because the model emits pixel coordinates and needs to know the coordinate space:

{
  "model": "claude-sonnet-4-6",
  "max_tokens": 4096,
  "tools": [
    {
      "type": "computer_20251124",
      "name": "computer",
      "display_width_px": 1280,
      "display_height_px": 800,
      "display_number": 1
    }
  ],
  "messages": [
    {"role": "user", "content": "Open the settings page and turn off email digests."}
  ]
}

The tool type is version-dated, and the version you use is tied to the model and to a beta header sent alongside the request. Both change across releases, and mixing a tool version with a model that does not support it is an error rather than a downgrade.

Computer use is a beta feature. The current tool type string, the required anthropic-beta header value and the per-model support matrix are all on Anthropic’s computer use documentation. Read them from there rather than from any page that pins a version string, including this one.

The round trip

Nothing happens on Anthropic’s side. The model emits an action; your harness performs it on a real screen; you send back what the screen now looks like. One full cycle:

// 1. The model asks for a screenshot.
{
  "content": [
    {"type": "text", "text": "Let me look at the current screen."},
    {"type": "tool_use", "id": "toolu_01C4…", "name": "computer",
     "input": {"action": "screenshot"}}
  ],
  "stop_reason": "tool_use"
}

// 2. You capture the screen and return it as an image block.
{
  "role": "user",
  "content": [
    {
      "type": "tool_result",
      "tool_use_id": "toolu_01C4…",
      "content": [
        {"type": "image",
         "source": {"type": "base64", "media_type": "image/png", "data": "iVBORw0KG…"}}
      ]
    }
  ]
}

// 3. The model reads the screenshot and emits a click, in pixels.
{
  "content": [
    {"type": "tool_use", "id": "toolu_01D8…", "name": "computer",
     "input": {"action": "left_click", "coordinate": [412, 336]}}
  ],
  "stop_reason": "tool_use"
}

// 4. You click, capture again, and return another screenshot. Repeat.

Two details make this different from an ordinary tool. First, tool_result.content is an array rather than a string, because it carries an image block — the array form of content exists precisely for this. Second, the loop is long: a task a person does in six clicks is a dozen or more round trips, each one carrying a fresh screenshot into the context.

The action vocabulary

Every call has an action discriminator in input, and the other fields depend on it. The families are:

  • Observation. screenshot takes no other fields and returns the current screen. cursor_position reports where the pointer is.
  • Pointer. left_click, right_click, middle_click, double_click and mouse_move take a coordinate pair. Drag actions take a start and an end.
  • Keyboard. type takes text and enters it as keystrokes; key takes a key name or chord for things a text string cannot express.
  • Scrolling. A scroll action takes a coordinate, a direction and an amount.

Coordinates are in the pixel space you declared in display_width_px and display_height_px. If your harness resizes the screenshot before sending it but keeps reporting the original geometry, every click lands in the wrong place — and the symptom is a model that appears to be guessing rather than a crash.

Resolution is a cost decision

Screenshots are images and images are tokens. A computer-use loop sends one on almost every turn, so the resolution you choose is multiplied by the length of the task. Higher resolution means the model can read smaller text and hit smaller targets; it also means more tokens per turn and a context window that fills faster.

The two levers are the resolution itself and how many screenshots you keep. Older screenshots in a long task are usually dead weight — nothing in them is still true — so pruning or truncating earlier image results is the single largest saving available in a computer-use loop. See how images are priced in tokens for the arithmetic.

Newer model generations changed the arithmetic by accepting higher-resolution screenshots, and with them a useful property: the coordinates the model emits map directly onto screenshot pixels rather than onto a downscaled space you have to convert back from. Where an older integration needed a scale factor between the screenshot it sent and the screen it clicked on, the current shape is one-to-one — and every scale-factor line left over from the older approach is now a bug that moves clicks by a predictable offset.

Where the loop breaks

The protocol is simple and the loop is not. Four failure modes account for most of the time lost, and none of them present as an API error.

  • Stale screenshots. Capture immediately after a click and you get the page mid-transition — a spinner, a half-painted modal, the old view still fading. The model reasons carefully about a screen that no longer exists and emits a click at a coordinate that was valid for a hundred milliseconds. The symptom is an agent that seems to be doing something inexplicable; the cause is that it is looking at the past.
  • Coordinate space drift. The declared display_width_px and display_height_px are what the model aims at. If your harness captures at one size and sends at another, or a window is not maximised, or a high-DPI display reports logical rather than physical pixels, every click lands offset by a constant. It looks like the model is bad at clicking, and it is actually arithmetic in your capture path.
  • Context exhaustion. Every turn adds a screenshot, and screenshots are the largest content type in the conversation. A long task fills the window with images of screens that are no longer on screen. This is the loop’s natural death, and it arrives as a context-window error rather than as a task failure.
  • The loop that will not end. If the element the model is looking for is not there — a dialog it expected, a button behind a scroll — it will keep looking. There is no stop condition in the protocol. Your iteration cap is the stop condition, and without one the loop runs until the context fills or the budget does.

There is one more shape worth recognising because it looks like a failure and is not. A long-running turn can come back with a stop_reason indicating the model paused rather than finished — a checkpoint in a long agentic run rather than an error. The correct handling is to send the conversation back to continue it, not to treat it as a truncation and retry from the top; retrying from the top repeats every action already taken, which on a computer-use agent means doing the work twice on a real machine.

What the harness owes you

The API gives you a model that names actions. Everything else is yours, and the parts that matter are not the fun ones:

  • Isolation. The model is acting on a real machine with whatever permissions that machine has. A container or VM with no credentials it does not need is the baseline, not a hardening step.
  • A settle delay. Screenshot immediately after a click and you capture the page mid-transition. The model then reasons about a screen that no longer exists. A short wait before capture removes a whole class of confusing behaviour.
  • A turn limit. The loop has no natural end if the model cannot find what it is looking for. Cap the iterations and surface the failure rather than letting it run.
  • A gate on irreversible actions. The model cannot tell a preview button from a submit button by anything except how they look. Anything that spends money, sends a message or deletes data should require a human, and that gate belongs in your handler rather than in the prompt.