Calibrating an Extraction Model's Confidence Scores Against Actual Error Rate
12 min read · updated August 11, 2026
Your pipeline reports a confidence of 0.90 on a field. Of all the fields it has ever reported at 0.90, what fraction were actually correct? If you cannot answer that with a number you measured, the score is an ordering, not a probability, and every threshold derived from it is derived from nothing.
Why the raw score is not a probability
Three different quantities get called “confidence” in extraction work, and none of them is P(correct) to begin with.
- A sequence probability. Exponentiating the summed log-probabilities over the value’s tokens tells you how likely that string was as a continuation, given the prompt and the image. That is a statement about the model’s own generative distribution, not about the page. A model that consistently misreads a smudged 8 as a 3 will emit the 3 with high probability every time.
- A classifier score. If a discriminative recogniser produced the value, its softmax output is a normalised score over classes. Modern deep networks are systematically overconfident here — Guo and co-authors documented the effect and its dependence on depth, width and weight decay in “On Calibration of Modern Neural Networks” (ICML 2017), alongside the temperature-scaling fix used below.
- A self-reported number. A
confidencekey the model filled in is generated like any other token. It is an ordinal signal at best; see scoring confidence per field for why its histogram spikes on round numbers.
Calibration is the step that converts any of these into the quantity you need. It does not make the model better at reading documents — a perfectly calibrated bad model is still bad. It makes the model honest about how bad it is, which is what lets a threshold and a queue do their jobs.
The reliability diagram
The measurement is simple and the discipline is in the details. Take a set of extracted fields for which you know the truth. Partition them into bins by reported confidence. For each bin, compute two numbers: the mean reported confidence in the bin, and the observed fraction correct. Plot observed against reported.
Perfect calibration is the diagonal. A curve that sits below the diagonal is an overconfident model: it says 0.9 and delivers 0.7, which is the usual direction and the dangerous one, because everything auto-accepted above your threshold is failing more often than the number implies. A curve above the diagonal is underconfident, which costs money rather than correctness — you are paying for reviews of fields that were fine.
Read the shape, not just the gap. A curve that is flat near 1.0 — reported 0.95 and reported 0.99 both delivering 0.86 — means the top of the score range carries no information, so no threshold placed inside that flat region can separate anything. That is a different disease from a uniformly shifted curve, and only one of them is fixed by rescaling.
Summarising it: ECE, MCE and Brier
Expected calibration error is the standard scalar summary: the weighted average of the gap between accuracy and confidence across bins, weighted by how many samples fall in each.
ECE = sum over bins k of (n_k / N) * | acc(B_k) - conf(B_k) | n_k samples in bin k N total samples acc(B_k) fraction of bin k that was correct conf(B_k) mean reported confidence in bin k MCE = max over bins k of | acc(B_k) - conf(B_k) | (worst bin) Brier = (1/N) * sum over i of (p_i - y_i)^2 (y_i is 1 or 0)
ECE is what you report. MCE is what you check before shipping, because a 0.02 ECE hides a bin at 0.95 that delivers 0.6 if that bin is small — and small high-confidence bins are exactly where auto-acceptance lives.
The Brier score, introduced by Glenn Brier in 1950 for weather forecasts, is the mean squared error between the predicted probability and the binary outcome. It is a proper scoring rule, meaning it is minimised by reporting your true belief, and unlike ECE it does not depend on a binning choice. Use it as the headline number when comparing two calibration methods.
Three ways the measurement lies
Bin count and bin width
ECE is not invariant to binning. Ten equal-width bins is the common default, and on extraction data it is a poor one, because reported confidences pile up above 0.9 — you end up with one enormous bin spanning 0.9 to 1.0 that averages away the whole region you care about, and nine nearly empty bins that dominate nothing. Use equal-mass bins (equal sample counts per bin, so the boundaries fall where the data is) and report the bin edges alongside the number. Then check that the answer does not move much when you change the bin count, because a figure that swings between 15 and 30 bins is an artefact.
Sample size per bin
Each bin is a binomial proportion, so its uncertainty is roughly 1.96 × sqrt(p(1−p)/n). At an observed accuracy of 0.95 with 100 samples in the bin, that is about ±0.043 — wide enough that the bin cannot distinguish a well-behaved 0.95 from a badly overconfident 0.91. To resolve differences of a percentage point in a high-confidence bin you need thousands of labelled fields in it, not hundreds. Plot the interval on the diagram; a reliability curve drawn without error bars invites people to interpret noise.
Selection bias in the labels
This is the one that quietly invalidates most in-production calibration work. Your ground truth comes from reviewed fields. If fields are routed to review because their confidence was low, then your labelled set is a sample of the low-confidence population, and the high-confidence bins are populated only by whatever happened to be reviewed for other reasons. Fitting a calibration curve on that sample produces a curve for a population you do not run.
The fix is structural: reserve a stratum of fields sampled uniformly at random, independent of confidence, and review those regardless. Two per cent of auto-accepted fields is often enough to anchor the top of the curve, and the general question of how much to sample is covered separately. When you combine the strata, weight each observation by the inverse of its probability of being reviewed, or you will re-import the bias you just paid to remove.
Fitting the correction
Once measured, the mapping from raw score to probability is fitted on held-out data. Three standard choices, in increasing order of flexibility and data appetite:
- Temperature scaling. One parameter. Divide the logits by a scalar
Tfitted to minimise negative log-likelihood on a validation set, then re-normalise. It cannot change the ranking of predictions at all, only their spread, which makes it the safest option and the one to try first. It needs the logits, so it is available when you own the recogniser and not when you only have a returned number. - Platt scaling. Fit a one-dimensional logistic regression from the score to the outcome — two parameters, a slope and an intercept. Works directly on a bare score, so this is the practical choice for a provider-returned confidence. Also monotone.
- Isotonic regression. Non-parametric, fits any monotone step function, and will match a curve that Platt cannot — including the flat-near-1.0 pathology. It is also the one that overfits on small samples, producing steps that are artefacts of a handful of observations. Do not reach for it below a few thousand labelled fields per field type.
Fit per field type, not once globally. The base rate and the error modes of a printed invoice number and a handwritten signature date have nothing to do with each other, and a single curve fitted across both is an average of two different functions that describes neither. Where a field type has too little data for its own fit, group it with fields that share a recogniser and a difficulty class, and say in the record which group was used.
Every fitted curve is stale the moment the pipeline changes. A new model version, a prompt edit, a preprocessing change, a different scanner in one office — each shifts the score distribution, and the curve fitted before it now silently mis-states the probability. This is why stamping the model and prompt version on every field is not paperwork: it is what lets you scope a calibration to the population it was fitted on.
The procedure
- Assemble a labelled set: extracted field, reported confidence, and the verified correct value. Include the random audit stratum, and record for each observation the probability that it was selected for review.
- Define “correct” before you start, in writing. Exact string match after normalisation is the honest default. Deciding case by case whether a near-miss counts is how a calibration study becomes unreproducible.
- Split by field type. Within each type, split the data into a fit set and an evaluation set — the curve must never be evaluated on the data it was fitted on.
- Bin the fit set by equal mass, compute per-bin accuracy and mean confidence, and plot the reliability diagram with binomial intervals. Look at it before computing any summary number.
- Fit Platt or temperature scaling on the fit set. Evaluate ECE, MCE and Brier on the evaluation set, both before and after. If the fitted curve does not improve the Brier score, it is not helping.
- Store the fitted parameters keyed by pipeline version and field type, and apply them at write time so that the number stored on the field record is already a probability. Downstream code should never have to know a raw score existed.
import numpy as np
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import brier_score_loss
def equal_mass_bins(p, n_bins=15):
edges = np.quantile(p, np.linspace(0, 1, n_bins + 1))
return np.unique(edges)
def ece(p, y, n_bins=15):
edges = equal_mass_bins(p, n_bins)
idx = np.clip(np.digitize(p, edges[1:-1]), 0, len(edges) - 2)
total = 0.0
for k in range(len(edges) - 1):
m = idx == k
if m.sum() == 0:
continue
total += m.mean() * abs(y[m].mean() - p[m].mean())
return total
# fit set / eval set, already split by field type
platt = LogisticRegression().fit(p_fit.reshape(-1, 1), y_fit)
p_cal = platt.predict_proba(p_eval.reshape(-1, 1))[:, 1]
print("ECE raw", ece(p_eval, y_eval), "-> cal", ece(p_cal, y_eval))
print("Brier raw", brier_score_loss(y_eval, p_eval),
"-> cal", brier_score_loss(y_eval, p_cal))When that runs and the calibrated Brier score is lower, the number on your field records means what it says, and a threshold derived from costs becomes a real decision rather than a habit.