Skip to content

Testing Correct Handling of a Provider's Announced Maintenance Window

9 min read · updated August 11, 2026

A provider posts a maintenance window for 02:00–04:00 UTC. Your error handling is good, your retries are tuned, and for two hours your application will retry an endpoint that is deliberately refusing, burning its retry budget on something no amount of retrying will fix.

Why maintenance is not just an error

A transient 500 and a two-hour maintenance window produce similar responses and demand opposite behaviour. Retrying a transient failure is right. Retrying a scheduled outage is a way of turning one provider’s planned work into your own outage, with your latency multiplied by your retry count and your queues backing up behind requests that cannot succeed.

The distinguishing signals are available and are usually ignored. A 503 with a retry-after header measured in thousands of seconds is not asking you to try again in a moment. Anthropic’s documented set includes a 529 overloaded_error for capacity, as distinct from a 500 api_error, and treating those two the same is the most common version of this mistake — the first should usually fail over, the second retry. And the provider’s status page or status API is a signal your application can consult, though you should treat it as advisory: status pages lag reality in both directions.

So the behaviour worth building, and therefore testing, is a state change rather than a per-request decision. When the evidence says “this provider is out for a while”, stop asking it, route elsewhere or degrade, and arrange to find out when it is back.

What to assert across the window

  • Attempts are bounded. Assert the total number of calls made to the unavailable provider over N user requests. If 500 user requests produce 1,500 provider calls, your retry policy is actively harming you. With a breaker open it should be a small constant plus the occasional probe.
  • The breaker opens, and quickly. Assert on the number of failures before the state changes, not on elapsed wall time, so the test is not timing-dependent.
  • Failover happens in the configured order. Assert which provider served the request, not merely that a request succeeded. A test that only checks success passes when failover picks the most expensive fallback first.
  • The degraded path is the one you designed. If there is no fallback provider, assert the specific degradation — a cached answer, a queued job with an accepted status, a defined error to the client — rather than an unhandled 500. Assert the status code and the body shape the client sees.
  • Latency during the window is bounded. A degraded response returned in 90 seconds is a timeout to the caller. Assert the fallback happens fast.
  • User-visible state is not corrupted. A conversation interrupted mid-window resumes correctly; a job does not lose its place. This is where a maintenance window causes lasting damage rather than temporary annoyance.

Simulating the window

Intercept at the HTTP boundary so the real client library, its own internal retries and its error parsing all participate. Mocking your own wrapper skips the layer where the behaviour actually lives — most SDKs retry transient failures internally with their own backoff, which means your call count is not what you think it is.

import { setupServer } from "msw/node";
import { http, HttpResponse } from "msw";

const maintenance = http.post("https://api.example.com/v1/messages", () =>
  HttpResponse.json(
    { type: "error", error: { type: "overloaded_error", message: "scheduled maintenance" } },
    { status: 529, headers: { "retry-after": "7200" } },
  ));

const fallbackOk = http.post("https://api.fallback.com/v1/messages", () =>
  HttpResponse.json({
    model: "fallback-model-1",
    content: [{ type: "text", text: "answer" }],
    stop_reason: "end_turn",
    usage: { input_tokens: 12, output_tokens: 4 },
  }));

export const server = setupServer(maintenance, fallbackOk);

Set onUnhandledRequest to fail when you start the server, so a request to an endpoint you did not model is a test failure rather than a silent pass-through to the network. During a failover test that is especially valuable: an unmodelled third provider quietly answering for real would make the test pass for entirely the wrong reason.

The test

import { describe, it, expect, beforeAll, afterEach, afterAll, vi } from "vitest";
import { server } from "./maintenance-server";
import { chat } from "../src/client";

beforeAll(() => server.listen({ onUnhandledRequest: "error" }));
afterEach(() => server.resetHandlers());
afterAll(() => server.close());

describe("provider maintenance window", () => {
  it("fails over in order and stops calling the unavailable provider", async () => {
    const primaryCalls = vi.fn();
    server.events.on("request:start", ({ request }) => {
      if (new URL(request.url).host === "api.example.com") primaryCalls();
    });

    const results = [];
    for (let i = 0; i < 50; i++) results.push(await chat("hello"));

    expect(results.every((r) => r.model === "fallback-model-1")).toBe(true);
    // Breaker opens after 5 consecutive failures; one probe per cool-down.
    expect(primaryCalls.mock.calls.length).toBeLessThanOrEqual(8);
  });

  it("honours a long retry-after rather than retrying immediately", async () => {
    const decision = classifyRetry({ status: 529, headers: { "retry-after": "7200" } });
    expect(decision.retrySameProvider).toBe(false);
    expect(decision.failover).toBe(true);
    expect(decision.cooldownMs).toBe(7_200_000);
  });
});

The second case is deliberately a unit test of the decision function rather than of the whole client. Retry policy is a pure function of the status, the headers and the attempt number, and testing it at that level lets you cover a dozen combinations in milliseconds instead of orchestrating a dozen server states. Keep the integration test for the wiring and the unit test for the matrix.

The recovery half

Most implementations of this are tested only on the way in, and the bugs are on the way out. Three cases are worth writing.

The breaker closes again. Restore the primary handler mid-test with server.use(), then assert that after the cool-down a probe succeeds, the breaker closes and subsequent requests go back to the primary. A breaker that opens correctly and never closes is a permanent failover nobody notices until the fallback bill arrives.

Recovery does not stampede. When the provider comes back, every queued request must not arrive at once. Assert that the probe is a single request and that traffic ramps rather than switching wholesale — the same coalescing problem as a cache stampede, with the same shape of fix.

Anything queued during the window drains correctly. If you queued jobs rather than failing them, assert they run once, in order where order matters, and that none was duplicated by a retry that fired while the provider was unreachable.

For the breaker’s own state machine and the thresholds worth setting, see circuit breakers; for the error taxonomy this depends on, the error-shape contract test.