Skip to content

Regression Testing a Summarisation Prompt Against Must-Keep Facts

10 min read · updated August 11, 2026

You changed four words in a summarisation prompt. The summaries still read well. Three weeks later someone notices that the dollar amount stopped appearing. A must-keep-facts test is the cheapest way to have found that on the pull request.

Why you cannot assert on the summary

The obvious test compares the new summary to a stored one. It fails on every run, because the model rephrases; and when you loosen it to an embedding-similarity threshold it stops failing at all, because a summary that drops one number is still 0.97 similar to one that keeps it. Both failure modes come from the same mistake: scoring the whole output when what you care about is a handful of specific propositions.

Invert it. Decide, per document, which facts the summary is required to contain, and assert each one separately. A dropped fact then produces a failure that says which fact, on which document, which is the difference between a red build you can act on and one you re-run.

This is the same move as a rubric, applied at the granularity where it can be automated: not “is this summary good” but “does the string 4.2 million appear, in a context that is not a negation”.

Building the must-keep list

Take five to twenty source documents that represent the shapes you care about — a short one, a very long one, one with a table, one with two conflicting numbers, one in another language if you serve one. For each, write between three and eight atomic facts. Atomic means one proposition: not “the merger details” but “the acquirer is Northwind” and “the price is $4.2bn” and “it closes in Q3 2027”.

Store them as data next to the document, not in the test file. The list is content, it will be edited by people who do not write TypeScript, and it wants to be diffable.

// fixtures/northwind.json
{
  "source": "northwind-8k.txt",
  "mustKeep": [
    { "label": "acquirer", "kind": "regex", "pattern": "Northwind" },
    { "label": "price", "kind": "number", "value": 4.2e9, "tolerance": 0 },
    { "label": "close date", "kind": "regex", "pattern": "Q3\\s*2027" },
    { "label": "regulatory condition", "kind": "claim",
      "text": "Completion is conditional on antitrust approval." }
  ],
  "mustNotContain": [
    { "label": "no invented advisor", "kind": "regex", "pattern": "Goldman|Morgan" }
  ]
}

The mustNotContain list is the half people forget, and it is where fabrication shows up. A prompt edit that makes summaries more fluent very often makes them more confident, and a bank that was never mentioned appears. One negative pattern per document costs nothing and catches a class of failure no positive check can.

Deterministic matchers first

Most facts can be checked without a model, and every fact that can be should be. A deterministic matcher is free, instant, and cannot itself regress.

  • Named entities — a case-insensitive regex, with the aliases you accept in the alternation. If the summary may legitimately say “Northwind Corp” or “the acquirer”, the pattern says so.
  • Numbers — normalise before comparing. Extract every numeric token from the summary, expand suffixes (bn, m, k), strip thousands separators and currency symbols, then look for a value within tolerance. Doing this as string matching is why “$4,200,000,000” fails a test that wanted “$4.2bn”.
  • Dates — same idea. Parse candidates out of the summary and compare parsed values, so a format change in the model’s output is not a regression.
  • Negation guard — a present token is not a present fact. “The price was not disclosed” contains neither the number nor a lie, but “Northwind did not acquire” contains the entity and inverts the claim. For entities that appear in a claim, check the sentence containing the match against a small negation pattern and fail if it fires.

When a matcher will not do

Some facts are genuinely propositional — “completion is conditional on antitrust approval” can be phrased fifty ways. For those, a model-graded check is reasonable, with three constraints that keep it from becoming the same unassertable mess you started with.

  1. Ask one binary question per fact, never “score this summary”. The prompt is the claim and the summary, and the required output is a structured object with a boolean and a quoted span from the summary that supports it.
  2. Require the span. A judge that must quote the supporting text cannot answer true for a fact that is absent without fabricating a quote you can then check with an exact substring match — which you do, in code, before trusting the boolean.
  3. Pin the judge. Fix its model identifier and its temperature at zero, and treat a change to either as a change to the test suite. A judge that silently upgrades makes every historical result incomparable, which is the failure mode described in silent model updates.

Judge calls cost money and time, so keep them to the facts that need them. In a twenty-document suite with six facts each, it is normal for a hundred of the hundred-and-twenty checks to be regex or numeric and twenty to be judged.

Sampling, thresholds and flakiness

One generation per document tells you about one sample from a distribution. At temperature zero that is nearly all there is to know; above zero it is not, and a test that runs a single sample will fail intermittently for reasons that have nothing to do with your edit.

Generate k summaries per document — three or five is usually enough to be useful — and assert on the retention rate per fact rather than on a single boolean. Then the suite has two kinds of failure and they mean different things:

  • A fact retained in 5 of 5 before and 0 of 5 after is a regression. The prompt edit removed it.
  • A fact retained in 5 of 5 before and 3 of 5 after is a stability regression, which is often worse: it will pass CI on the re-run and fail for a fraction of users forever.

Set the gate on the aggregate — for example, every fact labelled must-keep is retained in every sample, and any drop below that fails — and record the per-fact rates in the run output so the diff between two runs is readable. Store those rates alongside the prompt version, so the question “when did the price stop appearing” has an answer. That pairing of prompt version to per-fact retention is what turns this from a test into a quality regression signal.

The cost of this suite scales as documents × facts × samples, and the judged facts dominate it. Work out that number before you set k to ten: a hundred documents at five samples is five hundred summarisation calls per run, which is a per-commit bill somebody should have agreed to.