Detecting Sensor Calibration Drift Over Time
10 min read · updated August 11, 2026
A drifting sensor produces data that is smooth, plausible, in range, and wrong by an amount that grows. It is the hardest sensor fault to detect because the signature of drift and the signature of a real slow trend are, from inside the data, identical.
Drift is invisible from the data alone
Calibration drift is a slow change in the mapping from the physical quantity to the reported value — a gain error, an offset error, or both. Electrochemical gas sensors lose sensitivity as their electrolyte is consumed. Optical particulate sensors accumulate dust on the lens. Strain gauges creep. Reference voltages age. None of these produces a discontinuity, an out-of-range value, or an increase in noise.
The core epistemic problem: a series that rises by 0.4 units a month is consistent with the sensor drifting up by 0.4 units a month and with the measured quantity genuinely rising by 0.4 units a month. Nothing in the series distinguishes them. Detecting drift therefore always requires information from outside that series — a reference instrument, a peer sensor, a physical constraint, or a known condition. Any method that claims to detect calibration drift from one series in isolation is really detecting something else, usually a change in noise or an implausible rate of change.
That distinguishes it from the two neighbouring problems. A failing sensor changes its noise characteristics, which is detectable internally. Schema drift is a step change in units or encoding, usually at a firmware boundary. Calibration drift is a slow bias and needs an external anchor.
The periodic reference check
The standard anchor is a periodic comparison against something you trust more: a calibrated portable instrument, a laboratory assay of a grab sample, or a co-located reference-grade unit. The comparison gives you a difference at a known time, and a sequence of those differences is a series you can actually model.
Three requirements make this work, and each is commonly got wrong. The reference must be traceable and itself in calibration, or you are measuring the drift of the reference. The comparison must happen under matched conditions — same time, same location, same temperature — because a difference caused by the two instruments sampling different air is not drift. And the comparison must span the operating range: an offset error and a gain error look identical if you only ever check at one point, and they need different corrections.
Turning check readings into a drift rate
Suppose a CO2 sensor in a building is checked quarterly against a portable reference, both measuring the same outdoor air, and the differences are recorded in parts per million:
month sensor reference diff (ppm)
0 418 417 +1
3 429 421 +8
6 436 419 +17
9 447 423 +24
12 455 421 +34
Least squares fit of diff on month:
mean month = 6, mean diff = 16.8
Sxy = sum((m - 6) * (d - 16.8))
= (-6)(-15.8) + (-3)(-8.8) + 0(0.2) + 3(7.2) + 6(17.2)
= 94.8 + 26.4 + 0 + 21.6 + 103.2 = 246.0
Sxx = 36 + 9 + 0 + 9 + 36 = 90
slope = 246.0 / 90 = 2.73 ppm per month
intercept = 16.8 - 2.73*6 = 0.42 ppm
Residuals: -0.4, -0.6, +0.2, +0.4, +0.4 -> sd about 0.5 ppm
Standard error of slope = sd / sqrt(Sxx) = 0.5 / 9.49 = 0.053
95% interval on the slope is roughly 2.73 +/- 2*0.053 = 2.62 to 2.84The interval is what makes this a detection rather than an impression. A slope of 2.73 ppm per month with a standard error of 0.05 is unambiguously non-zero; five points with residuals of 20 ppm would have given the same slope and told you nothing. Reporting the drift rate without its uncertainty is the most common way a calibration programme produces confident nonsense, because the reference comparison itself is noisy and with three or four points that noise easily produces an apparent trend.
The intercept matters too. Here it is 0.42 ppm, so the sensor was essentially correct at installation and has drifted since. An intercept of 15 with the same slope would mean it was 15 ppm out of the box, which is a procurement problem rather than a drift problem, and the fixes are different.
When there is no reference: peer comparison
Most deployed fleets cannot afford quarterly manual checks on every unit. The substitute is redundancy: several sensors measuring quantities that should agree, or that should hold a fixed relationship. Four co-located air quality sensors on one rooftop, or twelve temperature probes in one cold store, give you a peer group.
The method is to monitor each unit’s deviation from the group median rather than from the group mean. The median is what makes it robust: with a mean, one badly drifted unit pulls the reference and smears blame across the whole group, whereas the median is unaffected until more than half the group has drifted. Fit a trend to each unit’s deviation series exactly as above.
The hole in this method is common-mode drift. If every unit is the same model from the same batch and they all age the same way, the group median drifts with them, every deviation stays near zero, and the fleet is uniformly wrong. Peer comparison detects relative drift only. The usual mitigation is a small number of reference-checked anchor units per site, deliberately from a different batch or technology, which turns the fleet into a hierarchy: anchors checked against a traceable instrument, everything else checked against the anchors.
Estimating bias as part of the model
Where drift is expected and slow, it can be estimated continuously rather than discovered quarterly, by making the bias part of the state being tracked. Augment a Kalman-style state vector with a bias term that has a very small process noise, so the filter is allowed to believe the bias moves, but only slowly:
state x = [ true_value , bias ]
model true_value_t = true_value_{t-1} + w1 (process noise Q1)
bias_t = bias_{t-1} + w2 (process noise Q2, tiny)
meas z_t = true_value_t + bias_t + v (noise R)
Q2 << Q1 is what separates them: fast changes are attributed to the
signal, slow persistent offsets accumulate into the bias term.This is only identifiable with an external constraint of some kind — a second unbiased sensor, or a known condition such as a period when the true value must be a specific number. Without one, the model has two unknowns and one equation, and the filter will happily distribute any offset arbitrarily between the two states. The same identifiability condition underlies every drift method: the information has to come from somewhere. The mechanics of the augmented state are in sensor fusion algorithms.
What to do once you have detected it
- Correct the stored data, but keep the raw. Apply the correction as a derived field with the calibration coefficients and their effective date recorded alongside. Overwriting raw readings makes the correction unauditable and irreversible when you later find the reference was itself out.
- Interpolating a correction backwards is a modelling choice. A linear interpolation between two calibration points assumes drift was linear in between, which is usually the best available assumption and is still an assumption. Record it as one.
- Retrain or re-baseline anything that learned the drifted data. An anomaly model fitted during a drifting period has absorbed the drift into its notion of normal, and correcting the data without refitting the model produces a burst of false alarms the day the correction lands.
- Let the drift rate drive the calibration interval. A unit drifting at 2.7 ppm a month against a 25 ppm tolerance has about nine months before it exceeds it. Scheduling by measured rate rather than by a fixed annual calendar sends technicians to the units that need them.
- Know when replacement beats recalibration. Sensors whose drift comes from consumption — electrochemical cells in particular — usually accelerate as they age, so a rate that is itself increasing is a sign the correction is about to stop working and the part needs replacing.