Skip to content

Using a Durable Object to Hold Agent State

10 min read · updated August 11, 2026

An agent needs to remember what it said. A Worker cannot, because it holds no state between requests. A Durable Object can, and the reason it can is worth understanding before you write one: it is a single instance, addressed by name, that all requests for that name are routed to.

Why a Durable Object and not KV

Conversation state has a property that eventually-consistent storage handles badly: it is read, modified and written on every single turn. Two requests for the same conversation arriving close together will both read the transcript, both append their turn, and both write — losing one. Eventually-consistent stores make this worse by also allowing the read to be stale.

A Durable Object removes the race by construction rather than by locking. There is exactly one instance per id, requests to it are handled one at a time, and Cloudflare documents input gates that prevent other events being delivered while a storage operation is in flight. You get read-modify-write correctness without writing a compare-and-swap loop, which is the actual reason to reach for one.

The class and its storage

A modern Durable Object extends DurableObject from cloudflare:workers and exposes ordinary methods, which the Worker calls over RPC on the stub. SQLite-backed objects give you ctx.storage.sql.exec(), and for a transcript that is a much better fit than a key-value blob.

import { DurableObject } from "cloudflare:workers";

type Turn = { role: "user" | "assistant"; content: string };

export class Conversation extends DurableObject<Env> {
  constructor(ctx: DurableObjectState, env: Env) {
    super(ctx, env);
    ctx.storage.sql.exec(
      "CREATE TABLE IF NOT EXISTS turns (" +
        "seq INTEGER PRIMARY KEY AUTOINCREMENT, " +
        "role TEXT NOT NULL, " +
        "content TEXT NOT NULL, " +
        "created_at INTEGER NOT NULL)"
    );
  }

  async append(turn: Turn): Promise<void> {
    this.ctx.storage.sql.exec(
      "INSERT INTO turns (role, content, created_at) VALUES (?, ?, ?)",
      turn.role,
      turn.content,
      Date.now()
    );
  }

  async recent(limit = 20): Promise<Turn[]> {
    const rows = this.ctx.storage.sql
      .exec("SELECT role, content FROM turns ORDER BY seq DESC LIMIT ?", limit)
      .toArray() as Turn[];
    return rows.reverse();
  }
}

Running the CREATE TABLE IF NOT EXISTS in the constructor is deliberate and not wasteful. The constructor runs when the object is brought into memory, which happens on first use and again after every eviction, so this is the only place guaranteed to run before any method. The statement is a no-op once the table exists.

Note ORDER BY seq DESC LIMIT ? followed by a reverse rather than ORDER BY seq ASC with an offset. You want the last N turns, and the number of turns is unbounded, so the query must be cheap at the tail rather than at the head.

Routing a request to the right instance

The Worker turns a conversation id into a stub. idFromName() is a deterministic hash, so the same string always resolves to the same object anywhere in the world.

// wrangler.jsonc
{
  "ai": { "binding": "AI" },
  "durable_objects": {
    "bindings": [{ "name": "CONVERSATION", "class_name": "Conversation" }]
  },
  "migrations": [
    { "tag": "v1", "new_sqlite_classes": ["Conversation"] }
  ]
}
export default {
  async fetch(request: Request, env: Env): Promise<Response> {
    const { conversationId, message } = await request.json<{
      conversationId: string;
      message: string;
    }>();

    const id = env.CONVERSATION.idFromName(conversationId);
    const convo = env.CONVERSATION.get(id);

    await convo.append({ role: "user", content: message });
    const history = await convo.recent(20);

    const result = await env.AI.run("@cf/meta/llama-3.1-8b-instruct", {
      messages: [
        { role: "system", content: "You are a terse assistant." },
        ...history,
      ],
    });

    await convo.append({ role: "assistant", content: result.response });
    return Response.json({ reply: result.response });
  },
} satisfies ExportedHandler<Env>;

The migrations block with new_sqlite_classes is not optional and is the step most often missed. Without it the class is not registered as a SQLite-backed Durable Object and ctx.storage.sql will not be available. Deploying a new class without a migration tag fails at deploy time, which is the good outcome; the bad one is registering it without SQLite and discovering it later.

Note also that idFromName takes whatever string you give it, so the conversation id is a capability. If it comes from a client without being checked against the authenticated user, anyone who guesses an id reads someone else’s transcript. Derive it from a server-side session, or namespace it with the user id before hashing.

Keeping the context window bounded

The recent(20) above is a crude bound and it will eventually be wrong in both directions — twenty short turns is too little context, twenty long ones may not fit the model. The better bound is on tokens, and the cheapest place to enforce it is in the object, because that is where the rows are.

Store an approximate token count per row at insert time, then select backwards accumulating until you cross a budget. It is approximate because you are not running the model’s tokeniser in a Worker, and approximate is fine if you leave headroom. The alternative pattern is to summarise: once the transcript passes a threshold, run a summarisation call, write the summary as a single row, and mark the turns it covers as superseded rather than deleting them, so you can audit what the summary replaced.

Deleting a conversation

Deleting one conversation is easy: add a method that calls this.ctx.storage.deleteAll() and the object’s storage is gone. Deleting all conversations older than ninety days is not, and the reason is a property of idFromName that is easy to miss.

  async purge(): Promise<void> {
    await this.ctx.storage.deleteAll();
  }

  async purgeOlderThan(cutoffMs: number): Promise<number> {
    const cursor = this.ctx.storage.sql.exec(
      "DELETE FROM turns WHERE created_at < ?",
      cutoffMs
    );
    return cursor.rowsWritten;
  }

idFromName is a deterministic hash, and it is one-way. Given a conversation id you can always find its object; given an object you cannot recover the conversation id that produced it. There is no list() over the names in a namespace, because the names were never stored — only their hashes were used to route.

The consequence is that a Durable Object namespace is addressable but not enumerable, and any retention policy, per-user export or “delete everything belonging to this account” request needs a second index that you maintain yourself. In practice that is a D1 or KV table mapping user id to conversation id with a last-active timestamp, written from the same Worker that creates the conversation. It is a small amount of bookkeeping and it is much cheaper to add on day one than to reconstruct on the day somebody asks for their data back.

The alternative is per-object expiry: have each object set an alarm well past its last use and call deleteAll() when it fires, rescheduling on every new turn. That gives you a TTL without a second index, and it gives you nothing for a targeted deletion request, so most systems that need both end up with both.

The ceilings you will actually meet

Cloudflare documents, for SQLite-backed Durable Objects at the time of writing: a combined key-and-value ceiling of 2 MB, 10 GB of storage per object on Workers Paid, and 500 Durable Object classes per account on Paid with an unlimited number of individual objects per class. It also documents a soft limit of about 1,000 requests per second to a single object.

Read those in order and the design falls out. The 2 MB ceiling is what punishes the naive design — one key holding the whole serialised transcript hits it and then the conversation simply stops working. Rows never hit it, because each row is one turn. The 10 GB figure means a per-conversation object will never fill up in practice. The 1,000 requests per second soft limit is the one that constrains the shape: objects scale by having many of them, not by making one fast, so an object keyed per conversation is fine and an object keyed per tenant is a bottleneck waiting to happen.

These are Cloudflare’s documented Durable Objects limits at the time of writing and differ between the Free and Paid plans; the Free plan figures are lower. Check Cloudflare’s Durable Objects limits page for current values.

The same class extends naturally to holding a live socket, which is what the WebSocket hibernation page builds, and the same single-threaded property is what makes a Durable Object a correct per-user rate limiter.