Tagging Regression Tests to a Specific Prompt Version
9 min read · updated August 11, 2026
A red case tells you an assertion failed. It does not tell you whether the case was written against the prompt you are editing right now or against one from four rewrites ago — and that is the difference between a regression and a case that was never going to pass again.
What a bare failure does not tell you
Consider two red builds that look identical in the terminal. In the first, a case written last week against the current prompt has started failing: something you just changed broke a behaviour that was working. That is a regression and you should not merge. In the second, a case written in March asserts that the assistant emits a refund_window_days field, and February’s rewrite removed that field from the response contract on purpose. That case is not reporting a regression; it is reporting that nobody deleted it.
Both print the same thing. The person triaging has to reconstruct the history by hand — open the case, blame the file, find the prompt commit nearest to it, read that revision of the prompt, decide. That is several minutes per case, on a budget that the sizing arithmetic shows is the binding constraint on how large the suite can be. Attach the answer to the case and the reconstruction disappears.
Version the content, not the file
The instinct is a version number in the prompt: v3, then v4. It fails the first time somebody fixes a typo without bumping it, and after that every tag in the suite is a lie you cannot detect. Use a hash of the template text, computed where the template is defined, so it cannot fail to change when the text does.
// prompts/refund.ts
import { createHash } from "node:crypto";
export const REFUND_SYSTEM = `You are a refunds assistant.
Reply with JSON matching the RefundDecision schema.`;
export const REFUND_ID = createHash("sha256")
.update(REFUND_SYSTEM)
.digest("hex")
.slice(0, 8); // e.g. "9f4c1ab2"Eight hex characters is plenty for a namespace of a few hundred revisions and short enough to sit in a test name. Hash the template only — not the rendered prompt with a fixture interpolated into it, or every case gets a different id and the tag stops naming a shared thing.
If your prompts are composed from fragments, hash the composed template, and hash each fragment separately as well. A change to a shared preamble then shows up as a changed id on every prompt that includes it, which is exactly the blast radius you want to see.
Three fields on every case
Keep the metadata with the case data, not in the test function name, so it is queryable without parsing strings.
// tests/regression/refund/cases.ts
export const CASES = [
{
caseId: "partial-refund-foreign-currency",
writtenAgainst: "9f4c1ab2",
mode: "fabrication",
origin: "ticket:SUP-4182",
input: { /* ... */ },
},
];writtenAgainst— the prompt id at the moment the case was written. It is written once and never updated automatically. A field that a script refreshes is a field that always matches, which tells you nothing.mode— the failure mode from your failure-mode taxonomy. This is what decides whether the case belongs on the gating tier.origin— where the case came from: a ticket id, an incident number, orhand. Six months later this is the difference between “a customer hit this” and “somebody imagined this”, and they are not deleted with equal enthusiasm.
Getting it into the failure output
Metadata nobody sees at the moment of failure is metadata that does not exist. The most portable place to put it is the test name, because every runner and every CI annotation prints that without configuration:
import { describe, test, expect } from "vitest";
import { REFUND_ID } from "../../prompts/refund";
import { CASES } from "./cases";
describe("refund regression", () => {
test.each(CASES)(
"[$mode] $caseId (written against $writtenAgainst)",
async (c) => {
const out = await runRefund(c.input);
expect(RefundDecision.safeParse(out).success).toBe(true);
},
);
});A failing line then reads [fabrication] partial-refund-foreign-currency (written against 9f4c1ab2), and if the current REFUND_ID is printed once in the run header, the comparison is a glance. Richer options exist — JUnit XML properties, a custom reporter that emits the fields as structured output — and they are worth it once you want to query across runs; check your runner’s reporter API for the current shape rather than assuming one. The test name costs nothing and works today.
The one number this unlocks
Here is the payoff that justifies the convention. At the start of every run, compare each case’s writtenAgainst against the current prompt id and count the mismatches:
const stale = CASES.filter((c) => c.writtenAgainst !== REFUND_ID);
console.log(
`prompt ${REFUND_ID}: ${stale.length}/${CASES.length} cases ` +
`were written against an earlier revision`,
);“83 of 120 cases were written against an earlier revision” is the single most useful line your suite can print, because it is the honest reading of what a green build meant. Those eighty-three cases passed, but they were written to describe behaviour the prompt no longer promises; their continuing to pass is at least partly luck. This is baseline drift made countable.
Print it, do not fail on it. A build that goes red because cases are old teaches people to edit the field, and a field people edit to make a build green is worse than no field. If you want a guard, put it on the ratio and set it loosely — a warning when more than eighty per cent of cases predate the current revision, which means it is time to read the suite rather than time to stop the merge.
Three tempting mistakes
Do not tag with a git commit SHA. It changes for reasons unrelated to the prompt — a rebase, a whitespace fix in a neighbouring file — so every case looks stale within a week and the signal is gone. The content hash changes when and only when the content does.
Do not bump the tag when you touch the case. writtenAgainst records when the assertion was reasoned about against a specific prompt. Fixing a typo in the fixture is not re-reasoning. Update it when you have actually re-read the prompt and confirmed the case still describes what it promises — and at that point you have done the work the field exists to prompt.
Do not put the tag only in the commit message. The information has to travel with the case into the failure output, and git history does not. The point of the convention is that the person triaging at 5pm never has to open a second tool.
One extension is worth the effort once the convention is in place: record the resolved model ID next to the prompt hash, in the run header rather than on each case. A case is written against a pair — a prompt revision and a model — and the second half moves without anyone editing a file. When a case that has been green for months goes red and both the prompt hash and the model ID in the header are unchanged from the last green run, you have narrowed the cause to the provider or the sampler before opening the diff, which is most of the work of investigating a silent model update.