Skip to content

Baseline Drift: Why Old Regression Tests Stop Meaning Anything

10 min read · updated August 11, 2026

A test that has been green for a year is usually read as a settled question. For a prompt suite it is more often a test that has stopped being connected to anything — and the three kinds of baseline come loose in three different ways, one of which is completely silent.

Three things people call a baseline

The word covers three quite different artefacts, and conflating them is why advice about baselines tends not to land.

  • A stored output. A snapshot: the text or the structure the model produced when somebody was satisfied with it, now committed and compared against on every run.
  • A stored threshold. A number: mean rubric score at least 0.82, exact-match accuracy at least 0.91, refusal rate below 3 per cent, measured once and encoded as a pass mark.
  • A stored judgement. A human decision recorded as a label — this output was acceptable, this one was not — reused later as ground truth for a case or as calibration for a model-judged score.

What moves underneath a fixed test

Baselines do not drift on their own. Four things move under them, and only two are yours.

The prompt moves, weekly in most teams. The model configuration moves: temperature, max tokens, tool schemas, retrieval corpus, the system preamble that some framework prepends. Those two you control and can at least correlate with a commit.

The model itself moves, and this is the one to be precise about because vendors differ. Anthropic’s models overview states that every Claude model ID is a pinned snapshot, that IDs containing a date are fixed to that release, and that from the Claude 4.6 generation onward the dateless IDs are also pinned snapshots rather than evergreen pointers — while for earlier generations the alias column entries are convenience pointers that resolve to a dated ID. Which is to say: whether the string in your config can start meaning a different model depends on which string it is, and you have to look it up rather than assume.

And the serving stack moves behind a fixed ID. OpenAI documents a system_fingerprint field on chat completions that “represents the backend configuration that the model runs with” and can be used with the seed parameter “to understand when backend changes have been made that might impact determinism”. The existence of that field is the vendor stating plainly that a pinned model ID is not a pinned computation.

Model naming policy, alias semantics and which fields a provider returns all change. Check the current model-versioning page for each provider you use rather than relying on the summary above.

Why stored outputs rot fastest

A stored full-text output is a baseline in which every character is an assertion, including the thousands of characters nobody meant to assert anything about. The model rewords a transition, reorders two bullets, or replaces a hyphen with an en dash, and the test goes red on a change that is not merely harmless but frequently an improvement.

What follows is the ritual that kills suites. Somebody runs the update flag — -u in both Vitest and Jest — the diff is enormous and mostly cosmetic, review scrolls past it, and the snapshots are accepted. That is fine the first three times. The fourth time, one of the accepted diffs is the total changing from 41.99 to 41.90, and it is accepted with the rest, because the reviewer’s job by then is to approve a wall of prose diffs rather than to read one.

The failure is structural, not a lapse in discipline. A baseline that goes red on harmless change trains the person reading it to stop reading, and a bulk-accept step is the mechanism through which that training becomes a shipped bug. It is why this cluster keeps returning to invariants: the point of asserting that the total equals the sum of the lines is that harmless rewordings do not touch it, so a red build means something and nobody develops the habit of clearing them.

If you keep snapshots at all, keep them of the structure rather than the prose — the set of keys, the field types, the tool names, the number of items — so that the snapshot only moves when the shape moves. That is a snapshot you can review honestly.

Thresholds fail in the other direction

Snapshot rot is loud. Threshold drift is not, and it is the more dangerous of the two precisely because nothing ever goes red.

You set a pass mark of 0.82 mean rubric score against the model you had in March. The provider ships a stronger model, or you migrate to one. Average quality rises to 0.91. The threshold now sits well below the floor of normal variation, so it will not trip — not for a prompt edit that degrades one behaviour, not for a retrieval bug that empties the context on ten per cent of requests, not for anything short of a catastrophe. The test is green, it will remain green, and it is measuring nothing.

The mechanism is that an absolute threshold encodes a relationship to a capability level that no longer exists. The fixes are all versions of making the comparison relative:

  • Record the achieved value on every run, not just the pass or fail, and watch the series. A drop from 0.91 to 0.86 is a signal even though both are over the mark.
  • Set the gate as a delta against the last known-good run on the same model ID, rather than as a constant. “No worse than 0.02 below the previous release” survives a model upgrade; “at least 0.82” does not.
  • Re-derive absolute marks when the pinned model ID changes. That change is the event; a calendar reminder is not.

Stored judgements decay the same silent way. A label recorded when the product answered support questions in one language is being reused as ground truth after the product started answering in six, and the labeller never considered the cases that now dominate. Judgement baselines need a re-labelling pass on a sample whenever the input distribution moves, which is a cost worth budgeting rather than discovering.

Keeping a baseline meaningful

Four habits, in the order they pay off.

Pin the model ID explicitly in the suite and print it in the run header along with the prompt hash, the SDK version and the sampling parameters. A run whose inputs are not recorded cannot be compared with any other run, which makes every baseline it produces unfalsifiable.

Treat a change of pinned ID as the trigger to revisit thresholds, not the calendar. Between ID changes the model is as fixed as the vendor can make it; at an ID change everything numeric in the suite is provisional again, which is also the argument for running the suite against both IDs in one job.

Convert prose baselines to invariant baselines wherever the invariant exists. Ask what must be true rather than what was said; the answer is usually shorter and it is always more durable.

And expire cases by prompt revision rather than by age. If you tag each case with the prompt hash it was written against, the suite can report how many cases predate the current revision. A case several revisions old is not automatically wrong, but it is automatically unexamined, and a suite that tells you how much of itself is unexamined is one you can still trust the green from.