Skip to content

Time-Aligning Data From Sensors With Different Clocks

10 min read · updated August 11, 2026

You plot two sensors that should respond to the same event and the traces are the same shape, offset horizontally. The correlation between them is far lower than it should be, and a model that uses both performs worse than one using either alone.

The symptom

The specific things people see, all of which are this problem:

  • A door-open event at 14:32:10 and the corresponding pressure transient at 14:32:12, consistently two seconds apart, in a system where the physical response is immediate.
  • A cross-correlation between two channels that peaks at a non-zero lag rather than at zero.
  • A joined table where the join on timestamp produces far fewer rows than expected, or a join with tolerance that matches the wrong rows.
  • Alerts firing in an order that contradicts causality — the downstream sensor reacting before the upstream one.
  • A fusion filter whose estimate is worse than either input, or an anomaly model whose residual is large and proportional to how fast the signal is moving.

That last one is diagnostic. If two series are misaligned by Δt, their difference is approximately the derivative of the signal times Δt. So the error is largest exactly when things are changing fastest, which is why this is so often first noticed as “the model only fails during transients”.

Four causes that look identical

  • Clock offset. The device’s real-time clock is simply set wrong, by seconds or by hours. Constant, and the easiest to fix.
  • Clock skew. The device’s oscillator runs slightly fast or slow. A typical uncompensated crystal is specified in the tens of parts per million; at 20 ppm a clock gains or loses about 1.7 seconds a day, so an offset corrected on Monday is back by Friday. Skew is also temperature dependent, so an outdoor device drifts differently in summer.
  • Timestamp applied at the wrong place. If the timestamp is written by the gateway or the ingest service on arrival rather than by the device at sampling, it includes transmission and queueing delay. Buffered devices are the pathological case: a reconnecting device flushes an hour of readings that all receive arrival timestamps within one second of each other.
  • Genuine physical lag. The pressure transient really does arrive two seconds after the door opens, because the air takes two seconds to move. Correcting this away is destroying real information, and it is the reason the diagnosis step below is not optional.

Telling them apart

  1. Check whether the lag is constant over time. Compute the cross-correlation lag on many separate days. A constant lag is offset or physical; a lag that grows linearly with time since last sync is skew, and the slope gives you the rate in ppm directly.
  2. Check whether it survives a device restart. Offsets from a clock that failed to sync usually reset; skew reappears immediately and grows again.
  3. Compare device timestamp to ingest timestamp. Store both. Their difference is the transport delay plus the clock error combined, and a histogram of it per device separates a stable clock error from variable network delay at a glance. If you are storing only one timestamp, this is the change to make first.
  4. Look for the flush signature. Many readings sharing a nearly identical ingest timestamp but spread device timestamps is a buffered device, and it means the ingest timestamp is unusable for that device.
  5. Rule out physics before correcting. Ask whether a lag of that magnitude is physically plausible for the coupling involved. If it is, the lag is a feature and the correct handling is to model it, not to remove it.

Estimating the offset from the data

Where you cannot fix the clocks retrospectively — and for historical data you never can — the offset can be recovered from the signals themselves whenever two channels share a common response.

import numpy as np

def estimate_offset(a, b, fs, max_lag_s=10.0):
    """Lag in seconds that best aligns b to a. Positive => b is late."""
    a = np.asarray(a, float); b = np.asarray(b, float)
    a = (a - a.mean()) / (a.std() + 1e-12)
    b = (b - b.mean()) / (b.std() + 1e-12)
    max_lag = int(max_lag_s * fs)
    corr = np.correlate(a, b, mode="full")
    lags = np.arange(-len(b) + 1, len(a))
    keep = np.abs(lags) <= max_lag
    corr, lags = corr[keep], lags[keep]
    k = int(np.argmax(corr))
    # parabolic interpolation for sub-sample resolution
    if 0 < k < len(corr) - 1:
        y0, y1, y2 = corr[k - 1], corr[k], corr[k + 1]
        delta = 0.5 * (y0 - y2) / (y0 - 2 * y1 + y2)
    else:
        delta = 0.0
    return (lags[k] + delta) / fs

Both series must be on a common uniform time base before this runs, which means resampling, and resampling assumes the timestamps within each series are internally consistent even if the series disagree with each other. Normalising each series first is what lets the correlation peak reflect shape rather than amplitude. The parabolic interpolation around the peak recovers sub-sample resolution, which matters when the offset is a fraction of the sampling interval.

Two conditions decide whether the estimate is trustworthy. There must be a shared transient in the window — two slowly varying signals with no common feature produce a flat correlation and a meaningless argmax, so check the peak is sharp and clearly above the surrounding values before using it. And periodic signals produce ambiguity at the period: two channels dominated by a 50 Hz component align equally well at every 20 ms shift, so the maximum lag bound must be smaller than the dominant period or the estimate will lock onto the wrong cycle.

The fix, and preventing a recurrence

  1. Correct historical data as a derived column, never in place. Add a corrected timestamp beside the original and record the offset applied and how it was determined. Overwriting makes the correction impossible to audit or reverse when the estimate turns out to be wrong.
  2. For skew, correct with a linear model, not a constant. Fit offset against time since last sync and apply the resulting rate. A single constant correction fixes the middle of the interval and makes both ends worse.
  3. Timestamp at the source, in UTC, at sampling. The device applies the time when it takes the reading, and the gateway adds its own arrival timestamp as a separate field. Keeping both is what makes every future diagnosis possible; local time and offsets are a presentation concern and belong nowhere near stored telemetry.
  4. Discipline the device clock continuously. On a network-connected device, NTP as specified in RFC 5905 estimates the offset from four timestamps — request sent T1, request received T2, reply sent T3, reply received T4 — as ((T2 - T1) + (T3 - T4)) / 2, with round-trip delay (T4 - T1) - (T3 - T2). The estimate assumes the path is symmetric, so an asymmetric route puts half the asymmetry straight into the offset. Over a typical internet path this lands in the milliseconds, which is ample for most telemetry and not for high-rate vibration work.
  5. Use hardware timestamping where milliseconds are not enough. IEEE 1588 Precision Time Protocol timestamps in the network interface rather than in software and, with supporting switches, reaches sub-microsecond agreement. It requires hardware support end to end, so it is a design decision rather than a retrofit.
  6. Monitor the offset as telemetry in its own right. Publish the device’s own estimate of its clock error and the time since its last successful sync, and alert on both. A device that has not synchronised for a week is producing data whose timestamps are quietly degrading, and this is the only way to know before the data is used.

Once alignment is trustworthy, the downstream work that depends on it — the windowing in streaming sensor data into a model and every cross-channel residual in correlated anomaly detection — becomes meaningful. Attempting either before fixing the clocks produces models that fit the misalignment.