Skip to content

Rate Limiting Requests Per User With a Durable Object

10 min read · updated August 11, 2026

Rate limiting is a read-modify-write on a shared counter, and that is exactly the operation distributed systems are worst at. A Durable Object is the one primitive on Workers that makes it exact, because the counter has only one owner and that owner does one thing at a time.

Why distributed counters are approximate

Consider a limit of 10 requests per minute per user, with the counter in an eventually-consistent store. Two requests arrive at two different data centres at the same moment. Both read 9. Both write 10. Both are allowed. The user made 11 requests and the counter says 10, and no amount of care in the calling code fixes it, because the race is in the storage semantics.

The usual mitigations are all approximations. Sharding the counter divides the limit and enforces each shard separately, so the real limit becomes a range. A compare-and-swap loop is exact but converts contention into retries, which is precisely wrong under the burst you are trying to limit. Local per-instance counters drift by however many instances you have.

A Durable Object sidesteps all of it. Cloudflare’s model guarantees one instance per id and single-threaded execution within it, with input gates preventing events from interleaving during a storage operation. Read-then-write is safe with no locking, no CAS and no retry, and the counter is the counter.

The limiter object

A token bucket is a better fit than a fixed window, because a fixed window permits double the limit across a boundary — ten requests at 11:59:59 and ten more at 12:00:00. The bucket refills continuously, so there is no boundary to exploit.

import { DurableObject } from "cloudflare:workers";

const CAPACITY = 10;          // burst size
const REFILL_PER_MS = 10 / 60_000; // 10 tokens per minute

type Bucket = { tokens: number; updatedAt: number };

export class UserLimiter extends DurableObject<Env> {
  async take(cost = 1): Promise<{ ok: boolean; retryAfterMs: number }> {
    const now = Date.now();
    const stored = (await this.ctx.storage.get<Bucket>("bucket")) ?? {
      tokens: CAPACITY,
      updatedAt: now,
    };

    const refilled = Math.min(
      CAPACITY,
      stored.tokens + (now - stored.updatedAt) * REFILL_PER_MS
    );

    if (refilled < cost) {
      const deficit = cost - refilled;
      await this.ctx.storage.put("bucket", { tokens: refilled, updatedAt: now });
      return { ok: false, retryAfterMs: Math.ceil(deficit / REFILL_PER_MS) };
    }

    await this.ctx.storage.put("bucket", {
      tokens: refilled - cost,
      updatedAt: now,
    });
    await this.ctx.storage.setAlarm(now + 10 * 60_000);
    return { ok: true, retryAfterMs: 0 };
  }
}

The cost parameter is the part that makes this worth building rather than using an off-the-shelf limiter. Model calls are not uniform: a 4,000-token request costs far more than a 200-token one, and a limiter that counts requests treats them identically. Charging tokens proportional to the expected work — or, after the fact, proportional to the usage.total_tokens the model reported — turns a request limiter into something much closer to a cost limiter.

Note that the read and the two possible writes happen with no await between the read and the decision. That is what makes the logic correct, and it is correct only because of the object’s single-threaded guarantee. The same code against a shared store would be the race described above.

Where the guarantee stops

The correctness above is not a property of Durable Objects in general. It is a property of that specific code, and it is easy to lose. What Cloudflare documents is input gates: while a storage operation is in flight, other events are not delivered to the object. That covers the read, the arithmetic and the write in take(), because there is nothing between them but computation.

Put an await on anything that is not storage into that window and the gate is not holding it. The object yields, another request is delivered, and you have rebuilt the race you came here to avoid:

  // BROKEN: the fetch reopens the window between read and write
  async take(cost = 1) {
    const stored = await this.ctx.storage.get<Bucket>("bucket");

    // a second request can be delivered during this await
    const plan = await fetch("https://billing.internal/plan").then((r) => r.json());

    await this.ctx.storage.put("bucket", {
      tokens: stored.tokens - cost * plan.multiplier,
      updatedAt: Date.now(),
    });
  }

The fix is the ordinary one for any critical section: do the slow, external work before you read, or after you write, but not between them. Fetch the plan first, then run the read-decide-write with nothing in the middle. If the external value is stable, cache it in a field and refresh it on a timer, accepting that hibernation and eviction will clear the field and the first request after a wake-up pays for the refresh.

Two further boundaries are worth stating plainly. The guarantee is per object id, so it holds for one user only if one user maps to exactly one id — a limiter keyed on an email address and a limiter keyed on a user id are two buckets for the same person. And the guarantee says nothing about availability: every limited request now depends on one object in one location, so it is a latency floor for every user and a failure domain for the ones whose object is unreachable. Decide in advance whether the limiter failing open or failing closed is the worse outcome, and write that decision down in the catch block rather than leaving it to whatever the runtime does.

Calling it from the Worker

export default {
  async fetch(request: Request, env: Env): Promise<Response> {
    const userId = await authenticate(request); // never trust a header alone
    const limiter = env.USER_LIMITER.get(env.USER_LIMITER.idFromName(userId));

    const { ok, retryAfterMs } = await limiter.take(1);
    if (!ok) {
      return new Response("rate limited", {
        status: 429,
        headers: { "retry-after": String(Math.ceil(retryAfterMs / 1000)) },
      });
    }

    const result = await env.AI.run("@cf/meta/llama-3.1-8b-instruct", {
      messages: [{ role: "user", content: await request.text() }],
    });
    return Response.json(result);
  },
} satisfies ExportedHandler<Env>;

Two details carry weight. The retry-after header is computed from the bucket rather than guessed, so a well-behaved client waits exactly long enough. And userId must come from something the client cannot set — deriving the object id from a request header means any caller can pick a fresh bucket by changing it, which is a rate limiter that does not limit.

The cost of this correctness is one extra round trip before every request, to whichever location the object lives in. For a user whose object was created near them this is small; for a user who has since moved continents it is not, and Cloudflare offers a location hint at object creation to influence where the instance is placed.

Cleaning up with an alarm

A per-user object for every user who has ever made a request is a lot of objects holding a tiny row each. The setAlarm call in take() schedules a wake-up, and the alarm handler deletes the state once the bucket has certainly refilled to full.

  async alarm(): Promise<void> {
    const stored = await this.ctx.storage.get<Bucket>("bucket");
    if (!stored) return;

    const full =
      stored.tokens + (Date.now() - stored.updatedAt) * REFILL_PER_MS >= CAPACITY;

    if (full) {
      await this.ctx.storage.deleteAll();
    } else {
      await this.ctx.storage.setAlarm(Date.now() + 10 * 60_000);
    }
  }

Deleting a full bucket is safe precisely because a missing bucket reads as full in take(). The two halves have to agree on that default or the cleanup becomes a way of granting extra quota.

When to use the built-in binding instead

Workers ships a rate-limiting binding, and for many cases it is the right answer with none of this code. It is configured declaratively:

// wrangler.jsonc
{
  "ratelimits": [
    {
      "name": "MY_RATE_LIMITER",
      "namespace_id": "1001",
      "simple": { "limit": 100, "period": 60 }
    }
  ]
}
const { success } = await env.MY_RATE_LIMITER.limit({ key: userId });
if (!success) return new Response("rate limited", { status: 429 });

Cloudflare documents period as accepting only 10 or 60 seconds, and the binding returns a single success boolean. That is the trade: it is fast, it costs no extra round trip to a specific object, and it gives you no retry-after, no variable cost per request and no window other than the two.

Choose the binding for coarse abuse protection. Choose a Durable Object when the limit is part of the product — a plan tier, a token budget, a per-customer quota you have to be able to explain on an invoice — because those need to be exact, need a cost dimension, and need to be readable. And note that neither of them protects you from the provider side: the ceilings described on the Workers AI rate limits page apply to your account as a whole, no matter how carefully you have shaped individual users.