Skip to content

A Rollback Plan for a Failed Re-Embedding Migration

10 min read · updated August 11, 2026

A re-embed you cannot undo is not a migration, it is a one-way door with a large bill on the other side. The plan below costs one extra write path and some disk, and it turns a bad outcome into a thirty-second revert.

The shape of the plan

Three commitments make everything else follow. First, never overwrite the live vector column in place: the new embeddings go into a new column, table, collection or namespace while the old one keeps serving. That is required anyway, because a half-converted index returns confident nonsense, but it is also what makes a rollback exist at all.

Second, the read path chooses its index from configuration, not from code. If the query builder names chunks.embedding_v2 as a literal, reverting means a deploy; if it reads RETRIEVAL_INDEX=chunks_v2 from config, reverting means changing a string. Those are different orders of magnitude of response time during an incident, and the difference costs you one indirection.

Third — and this is the one people leave until it is too late — the ingest path writes to both generations for the whole duration, starting before the backfill does. Without it, the old index stops receiving new documents the moment the backfill begins, so by the time you want to roll back, your rollback target is missing everything ingested since. A stale index is not a rollback target; it is a different outage.

The procedure

  1. Build the retrieval evaluation set first, before anything else. Fifty to two hundred real queries with the chunk ids that should be retrieved for each, drawn from production traffic and redacted. You cannot judge the new model without it and you cannot build it after the fact, because the old index will be gone. If this step is skipped the rest of the plan is theatre.
  2. Measure the old index against that set. Recall at your production k, the fraction of queries returning nothing above threshold, and the score distribution. This is the baseline every later decision compares against.
  3. Create the v2 destination. New column or new table with the correct width, no index yet. Decide the index type and any quantisation now, because it constrains the width you write.
  4. Deploy dual-writes and let them run for a day. Every create, update and delete in the ingest path writes to both v1 and v2. Deploy this on its own, before the backfill, and confirm it is working — a dual-write that silently drops v2 writes is invisible until the backfill finishes and the row counts disagree.
  5. Backfill in resumable batches. Order by primary key, keep a watermark of the highest id completed, commit the watermark in the same transaction as the rows. A crashed worker resumes; a re-run is a no-op. Rate-limit yourself to the ceiling derived in the cost and duration arithmetic rather than discovering it through 429s.
  6. Reconcile before you build the index. Count rows with a null v2 vector; it must be zero. Then check for rows written by the dual-write during the backfill that the backfill also touched, and confirm the later write won.
  7. Build the v2 index concurrently, then recalibrate the threshold against the evaluation set, as below.
  8. Flip a fraction of read traffic by a stable hash of tenant or session id — never per request, or one user’s follow-up question retrieves from a different index than their first and produces a bug report nobody can reproduce. Hold at each step long enough to see your slowest signal.
  9. Flip fully, keep dual-writes running. The rollback stays available for as long as v1 is being written to, and no longer.
  10. Stop dual-writing and drop v1 only when the exit conditions below are met.

The threshold does not travel

This deserves its own section because it causes more false rollbacks than genuine model regressions do. Two embedding models produce different score distributions for the same query set. A cutoff of 0.75 that gave you a median of six results per query on v1 might give you two on v2, or eighteen. Neither is a quality change; it is a calibration difference, and transplanting the constant is what makes it look like a collapse.

Recalibrate by matching a distribution rather than a number. Run your evaluation queries against v2, collect the scores, and choose the threshold at the quantile that reproduces the old median result count. Then evaluate quality at that threshold. Concretely: if 0.75 on v1 kept the top 12% of candidate scores, find the v2 score at that same percentile and start there.

# per query, on each index, over the evaluation set
scores_v1 = [s for q in queries for s in top_scores(v1, q, k=50)]
scores_v2 = [s for q in queries for s in top_scores(v2, q, k=50)]

keep = fraction_above(scores_v1, threshold_v1)   # e.g. 0.12
threshold_v2 = quantile(scores_v2, 1 - keep)

# only now compare recall@k on v1 at threshold_v1
# against recall@k on v2 at threshold_v2

The same warning applies to anything else tuned against score magnitudes: a deduplication cutoff, a semantic cache’s hit criterion, a “no good answer, decline to respond” guard. Each is a constant fitted to a distribution that no longer exists, and each will change behaviour at cutover. Find them before you flip; testing for similarity threshold drift covers how to catch them automatically.

Making the rollback trigger measurable

“Roll back if it looks bad” is not a trigger, because during a migration everything looks slightly bad and nobody wants to be the person who called it. Write the trigger down before the flip, in numbers the system already produces:

  • Empty-result rate. The fraction of queries returning nothing above the recalibrated threshold. Compare against the v1 baseline, not against zero.
  • Recall on the evaluation set. The one direct measure of retrieval quality you have. A defined drop against baseline is an immediate revert.
  • Answer-level citation rate. The fraction of generated answers that cite at least one retrieved chunk. This catches the case where retrieval returns something and the generator correctly declines to use it.
  • A structural check on citations. Whether the cited chunk actually contains the asserted claim, checked by string or entity overlap. Crude, cheap, and the fastest proxy for the slow signal of a user noticing.
  • Latency at the tail. A new index type or a wider vector changes p99, and a retrieval step that quietly went from 40ms to 400ms is a regression even if quality held.

Give each a threshold and a person. The trigger is not a discussion; it is a condition that, when met, causes the config flip without a meeting.

Deleting the old index, and when you may not

The plan needs an end, or the dual-write path becomes permanent architecture that nobody remembers the reason for. Drop v1 when three things are true: the rollback trigger has stayed quiet for a window longer than your slowest feedback loop, the evaluation set says v2 is at least at parity, and the storage cost of holding both has become worth reclaiming.

There is one condition that overrides all three, and it is why this page sits in this cluster. A rollback to v1 is only possible while the old embedding model still exists to serve queries — the query has to be embedded by the same model as the documents, so if the old model is being retired, the rollback expires on its retirement date whatever your dual-write does. If that is your situation, the migration has a hard deadline and the rollback window ends before it, not on it. Work backwards from the retirement date on your calendar and subtract the evaluation cycle, the ramp and one slipped sprint. Starting a re-embed a fortnight before the old model dies means running it without a rollback, which is the situation this entire page exists to avoid.