Skip to content

Prioritizing a Human Review Queue by Financial Impact, Not Arrival Order

10 min read · updated August 11, 2026

A review queue in arrival order treats a €4 stationery invoice and a €480,000 construction payment as equally urgent. If your reviewers clear the queue every day this costs nothing. The moment they do not, arrival order is a decision about which errors escape, made by whichever supplier’s post arrived first.

What FIFO is optimising

First-in-first-out is not arbitrary — it optimises fairness in waiting time, which is the right objective for a queue of people. For a queue of fields it optimises nothing anybody wants. It is also the ordering that makes a backlog invisible: every item is eventually reached, so the queue looks like it is working right up until the day it is 40,000 items deep and the oldest untouched item is three weeks old and attached to a payment that already went out.

The distinction to hold on to is that this is a different question from the one a threshold answers. A confidence threshold decides whether reviewing a field is worth its own cost — a question about one field in isolation, with an answer that does not depend on how many other fields exist. Ordering only arises because reviewer time is finite: the threshold has already declared these fields eligible, and there are more of them than there are minutes.

Deriving the ranking

Set it up as a budget problem. You have T reviewer-minutes in a shift. Each eligible field i takes t_i minutes to review and, if reviewed, avoids an expected loss:

v_i = P(field i is wrong) * r * impact_i
    = (1 - p_i) * r * impact_i

  p_i      calibrated probability the extracted value is correct
  r        probability the reviewer catches an error that is present
  impact_i cost if this particular wrong value escapes

Choosing a subset of fields to fit in T minutes while maximising total v is a knapsack problem. Its fractional relaxation has a known optimum — take items in decreasing order of value per unit of cost — and since one review is small relative to a shift, the greedy ratio ordering is very close to optimal here. So the priority score is:

priority_i = v_i / t_i = (1 - p_i) * impact_i / t_i

(r is constant across items and drops out of the ordering)

Two things fall out of the derivation that the usual folklore version — “sort by amount times confidence gap” — leaves out.

The first is the t_i denominator. A field that takes eight minutes to verify because it needs a lookup in another system has to clear a bar eight times higher than one that takes a minute, and ignoring that is how a queue ends up spending a whole shift on four items. If you have per-field-type handling times, use them; if not, a coarse two- or three-tier estimate already beats a constant.

The second is that 1 − p_i must be a calibrated probability for the products to be comparable across fields. If it is a raw score, then a field type whose scores happen to sit lower gets systematically prioritised regardless of whether it is actually less reliable, and your queue is sorted by an artefact of the scoring distribution. This is the dependency that makes calibration load-bearing rather than academic.

A worked queue

Five eligible fields arrive in this order. Impact figures are labelled assumptions; handling times are assumed at two tiers, one minute for a field with a source crop and four minutes for one requiring an external lookup.

#  field                 impact   p     1-p    t   v = (1-p)*impact   v/t
1  ref number             40      0.72  0.28   1        11.2            11.2
2  line description        8      0.55  0.45   1         3.6             3.6
3  payee account      120000      0.981 0.019  4      2280.0           570.0
4  invoice total       18400      0.93  0.07   1      1288.0          1288.0
5  vendor VAT id        2500      0.88  0.12   4       300.0            75.0

FIFO order,   first 4 minutes reviewed: #1, #2, #3(partial)  -> ~14.8 avoided
Ranked order, first 4 minutes reviewed: #4, #3              -> ~3568 avoided

The gap is not a tuning improvement, it is two orders of magnitude, and it comes almost entirely from one item. That is the usual shape: expected loss in a document queue is heavily concentrated, so most of the available benefit is captured by getting the top few items right and very little of it by refining the ordering of the tail.

Note also item 3 versus item 4. The account number has much higher impact but much higher confidence and a longer handling time, and the ratio puts the invoice total first. A ranking on impact alone would get that backwards; a ranking on confidence alone would put the ref number first and waste the shift.

Impact is not always an amount

“Financial impact” is a good default and a bad universal rule, because plenty of high-consequence fields carry no number at all. Make impact_i a function of the field and the record rather than a column, and let it combine:

  • The monetary amount on the record, where there is one — and use the document total rather than the individual field for fields like a bank account or a payment date, since a wrong account misdirects the whole amount, not its own value.
  • A fixed severity for identity and routing fields. A wrong patient identifier, a wrong tax identification number or a wrong counterparty is a serious event at any amount, and deserves a floor that keeps it out of the tail.
  • Regulatory or contractual exposure, which is usually a step function rather than a linear one — above a reporting threshold, the cost of a wrong figure jumps.
  • Irreversibility. A field feeding an outbound payment that executes tonight is worth far more review than the same field on a record that will be reconciled next month. Time-to-effect belongs in the impact function, and it is the single most useful non-monetary term.

Keep the function in one place, version it, and store the impact value that was used on the queue item. When somebody asks in six months why a particular field was never reviewed, the answer has to be reconstructible, and it is not if the priority was computed on the fly from inputs that have since changed.

Ageing, deadlines and half-done documents

A pure ratio ordering starves the bottom of the queue permanently. Sometimes that is correct — a field whose expected loss is €0.30 genuinely should never be reviewed, and it should be closed out rather than left to rot. But two cases need explicit handling.

Hard deadlines. Some items stop being useful at a known time: a payment run at 16:00, a customs filing, a month-end close. Deadline handling does not belong in the ratio; promote items whose deadline is inside the next scheduling horizon to the front outright, and let the ratio order what remains. Mixing a deadline into a score as a weight produces items that miss their deadline by a little and nobody can explain why.

Partially reviewed documents. If three fields on one invoice are eligible and the ranking scatters them across three shifts, you pay the cost of loading that document into a reviewer’s head three times, and the document is blocked the whole while. Group by document and rank the groups, or at minimum give a strong bonus to fields on a document that already has one field in progress. The setup cost of a document is real and it is not in the per-field t_i.

A light ageing term on top of the ratio is worth having as a safety valve even if you believe the tail should starve, because it converts a silent permanent backlog into something that drains. What matters more is reporting the tail: the total expected loss of eligible fields that were never reviewed is the number that tells you whether to hire a reviewer or improve the model, and it is only computable if you kept the priorities.

Ranking blinds your measurement

This is the consequence people discover late, and it is worth as much attention as the ordering itself.

Once the queue is ranked, the fields a human ever looks at are overwhelmingly high-impact and low-confidence. Those reviews are your only source of ground truth. So the labelled data your accuracy metrics, your calibration curve and your error analysis are all built from is a sample drawn with a probability that depends on exactly the variables you are trying to study. Accuracy measured on it will look terrible, calibration fitted on it will be wrong in the region where auto-acceptance happens, and neither error is visible from inside the data.

Two compensations, and you want both:

  • A random audit stratum. A fixed small percentage of all extracted fields — including the ones that were never eligible for review — sampled uniformly and reviewed regardless of priority. This is the only unbiased view of the population you actually ship. Budget for it explicitly, as a slice of reviewer capacity, or it will be the first thing dropped in a busy week.
  • Inverse-probability weighting. When you combine audit-stratum reviews with priority-driven reviews for any statistic, weight each observation by the reciprocal of its selection probability. Otherwise you have simply re-imported the bias into a larger sample. Store the selection probability on the queue item at the time it was chosen; reconstructing it afterwards from a priority formula that has since been edited is not possible.

The sampling rate itself is a separate trade-off, covered in choosing a human review sampling rate. The point here is narrower and easy to miss: prioritisation is not only a scheduling decision, it is a decision about what your organisation will be able to know about its own error rate.