Skip to content

Building a Review Queue That Shows the Source Location for Each Extracted Field

11 min read · updated August 11, 2026

A reviewer is shown a field name, an extracted value and a 14-page PDF. Almost all of the time they spend on that item goes into finding where the value came from. Everything in this page is about deleting that step, and the obstacle is that the component which produced the value usually has no idea where it was.

Why locating dominates review time

Decompose a field review into three parts: find the field on the page, read it, decide whether the extracted value matches. The second and third are seconds of work. The first is the whole job on a multi-page document, and it scales with document length while the other two do not.

That asymmetry is why source highlighting is the highest-leverage thing in a review interface, and it feeds straight back into the economics: C_review in the threshold derivation is proportional to handling time, so halving handling time approximately halves the confidence at which review stops being worth it — meaning you can afford to check more fields with the same headcount. Source highlighting is not a UI nicety; it changes what your pipeline can safely automate.

A second, less obvious benefit: a crop is auditable. A reviewer who approved a value while looking at a specific region of a specific page has left a much stronger record than one who approved it while looking at a document, and that record is what the field-level audit trail stores.

Three coordinate systems, one page

Getting a box onto the screen in the right place means moving between three spaces, and mixing them up produces boxes that are systematically flipped, offset or scaled — a bug that looks like a bad model and is not.

  • PDF user space. Defined by ISO 32000, the PDF specification maintained through ISO and documented publicly by the PDF Association. Its origin is the bottom-left of the page and its default unit is 1/72 inch, so y increases upwards. The visible page is the CropBox where one is present, not the MediaBox, and neither is guaranteed to start at zero — a page can legitimately have a box whose lower-left corner is at a non-zero offset, and boxes drawn without subtracting it are shifted for every field on the page.
  • Raster image space. What OCR sees after rasterisation, and what a scan is natively. Origin at the top-left, unit is the pixel, y increases downwards. The conversion depends on the render resolution: px = pt × dpi / 72, and the vertical axis flips, so y_img = (pageHeight_pt − y_pdf) × dpi / 72. Half of all wrong-box bugs are a missing flip and the other half are a wrong DPI.
  • Viewport space. Wherever the reviewer’s browser has drawn the page, at whatever zoom. Never store this. Store normalised coordinates in [0, 1] relative to the visible page box and let the front end multiply by its rendered size; then a change of zoom, device pixel ratio or thumbnail size cannot desync the overlay.

Page rotation is the fourth trap. A /Rotate 90 entry means the page is displayed rotated relative to its content coordinates, so a box that is correct in user space appears on the wrong edge. Decide once whether your stored coordinates are pre- or post-rotation, record which in the field record, and apply it in exactly one place.

The hard part: the model does not give you a box

Ask a general vision model for bounding boxes alongside the values and you will get plausible-looking numbers. They are generated tokens. Some will be approximately right, some will be off by a quarter of the page, and nothing in the response distinguishes the two. Highlighting the wrong region is worse than highlighting nothing, because a reviewer who is shown a box trusts it and checks the wrong thing.

The reliable construction runs the other way. Get coordinates from a component that measures them — an OCR engine over the rasterised page, or the text layer of a digital PDF — and then relocate the model’s extracted value inside that coordinate-bearing word stream by matching the string. OCR output formats are built for exactly this: hOCR carries a bbox per word in its title attribute, and ALTO, maintained by the Library of Congress, carries HPOS, VPOS, WIDTH and HEIGHT per String element.

And here is where it breaks, reliably, in a way that is worth designing for up front: the value in your record has usually been normalised, and the string on the page has not. The page says £1.234,56 and your record says 1234.56. The page says 3 Feb 26 and your record says 2026-02-03. The page hyphenates an identifier across a line break. Exact string search finds none of these, and a fuzzy search that is loose enough to find them is loose enough to match the wrong number elsewhere on the page.

The fix is to make the model return both forms. Ask for the normalised value and the verbatim substring exactly as printed, character for character, including its currency symbol, separators and casing. Match on the verbatim form; store the normalised form as the value. This costs a few output tokens per field and removes the entire class of failure.

{
  "total_amount": {
    "value": 1234.56,
    "verbatim": "£1.234,56",
    "status": "present",
    "source": {
      "page": 3,
      "rotationApplied": true,
      "bbox": [0.712, 0.184, 0.869, 0.203],
      "matchQuality": "exact"
    }
  }
}

matchQuality earns its place. Record whether the relocation was exact, fuzzy or failed, and let the interface degrade honestly: an exact match gets a solid highlight, a fuzzy match gets a dashed one and a tooltip, and a failure gets no box at all rather than a guess. A reviewer can work with “we could not locate this”. They cannot work with a confident box over the wrong figure.

When the value appears four times

An invoice total appears in the summary block, in the payment instructions, in the remittance slip and sometimes in a footer on every page. String matching finds all of them, and picking the first is arbitrary.

Rank the candidates rather than picking one. Three signals do most of the work: proximity to a label token that matches the field (“Total”, “Amount Due”, “Balance”) measured in the same coordinate space; agreement with the reading order the document actually has, which on a multi-column layout is a clustering problem rather than a text-order problem; and consistency with the boxes already assigned to neighbouring fields, since fields that belong to one block are usually spatially adjacent. Keep the runner-up candidates in the record. A reviewer who sees the highlighted region is wrong can then cycle to the next candidate instead of abandoning the tool.

The multi-occurrence case is also a free validation signal. If a value appears four times on the page and all four readings agree, that is independent corroboration of a kind a confidence score cannot give you. If three agree and one differs, you have found either an OCR error or a document that contradicts itself, and both are worth surfacing.

What to store, and in what units

  • Normalised coordinates, in [0, 1] relative to the visible page box, as [x0, y0, x1, y1] with a documented origin. Resolution-independent, so re-rendering the page at a different DPI for a different screen cannot break the overlay.
  • A page index, and be explicit about whether it is zero- or one-based. This is a trivial thing to get wrong and it puts every highlight on the wrong page.
  • The verbatim string, always. Even when the box is missing, a verbatim substring lets the interface search the page and lets a later investigation see what was actually printed rather than what your normaliser produced.
  • A quad rather than a rectangle where scans are skewed. Four corner points survive a two-degree rotation that an axis-aligned rectangle turns into a box with white margins and clipped ascenders.
  • A document and page fingerprint — a hash of the rendered page — so that if the source file is ever re-processed or re-rendered, a stale coordinate cannot silently point somewhere else on a page that has changed.

Wiring it up

  1. Rasterise each page once at a fixed, recorded DPI, and keep that number with the document. Every coordinate conversion afterwards depends on it, and a pipeline that renders at 200 dpi in one service and 300 in another will produce boxes that are correct in one and 50% off in the other.
  2. Run OCR to obtain a word stream with per-word boxes, even if a vision model is doing the actual extraction. The OCR pass exists here as a coordinate source, not as a competing extractor.
  3. Extract with a schema that requires both value and verbatim for every field.
  4. Relocate: normalise whitespace, search the word stream for the verbatim string allowing word-boundary merges and hyphenation at line ends, rank candidates by label proximity, and record matchQuality.
  5. Convert to normalised page coordinates, applying the crop-box offset and the rotation exactly once, and store alongside the field.
  6. In the interface, render the page image and draw the box from the normalised coordinates times the rendered size. Open the reviewer directly on the crop, with the full page one keystroke away.

The last step is where the time is actually saved, and it is worth being deliberate: default to the crop, not to the page. A reviewer who has to zoom to the highlight on every item is paying a smaller version of the same locating cost you built all of this to remove.