Skip to content

Compressing Sensor Data Streams for Storage

10 min read · updated August 11, 2026

Time-series telemetry compresses far better than general data, for a structural reason: consecutive timestamps are nearly evenly spaced and consecutive values are nearly equal. Encoders that exploit those two facts beat general-purpose compressors by a wide margin at a fraction of the CPU.

Why generic compressors underperform here

A naive representation of a sensor point is a 64-bit timestamp and a 64-bit float: 16 bytes. Handed to gzip, this does moderately well on the timestamps, which are highly regular, and poorly on the floats, because the low mantissa bits of a physical measurement are close to random and a byte-oriented dictionary coder finds no repeated substrings in them.

The insight that time-series databases are built on is that the redundancy is arithmetic, not textual. Timestamps at a fixed interval have a constant first difference. Values from a physical process change slowly, so consecutive floats share most of their high bits. Neither fact is visible to a general compressor operating on bytes, and both are trivial to exploit if you know the data is a series.

Delta-of-delta and XOR

The reference description is Facebook’s Gorilla paper, presented at VLDB 2015 as “Gorilla: A Fast, Scalable, In-Memory Time Series Database”, and the scheme it describes is now the basis of several open implementations. Two separate encoders:

  • Timestamps: delta of delta. Store the first timestamp in full, then the difference between successive timestamps, then the difference between successive differences. On a stream sampled at an exact interval, that second difference is zero, and zero is encoded in a single bit. A stream with mild jitter produces small integers that fit in a handful of bits with a variable-length code.
  • Values: XOR against the previous value. XOR two IEEE 754 doubles that are close in magnitude and the result is mostly leading zeros, because the sign, exponent and high mantissa bits are identical. Encode the number of leading and trailing zero bytes and only the meaningful bits between them. An unchanged value XORs to zero, again a single bit.

The paper reports that this combination brought the average storage per data point down to 1.37 bytes from 16, across Facebook’s production monitoring workload. That is a measurement of their data, not a guarantee about yours: the ratio depends entirely on how regular your timestamps are and how slowly your values move. A noisy 24-bit ADC sampled irregularly compresses much worse than a temperature reading that changes in the third decimal place.

For a second published data point on the same mechanism, the Prometheus storage documentation states that on average Prometheus uses only around one to two bytes per sample, which is the same order and arrived at the same way.

Both figures are vendor-published characterisations of specific workloads and are the kind of number that gets revised as implementations change. Verify against your own data before sizing anything on them.

A year of storage, derived

The inputs below are assumptions, chosen to be plausible for a mid-size industrial deployment. Substitute your own; the arithmetic is the part worth keeping.

ASSUMPTIONS
  channels                 200 sensors
  sampling rate            10 Hz per sensor
  raw point                16 bytes (int64 timestamp + float64 value)
  lossless ratio           1.37 bytes/point  (Gorilla paper, VLDB 2015)
  retention                12 months

POINTS
  per second   200 * 10                     =        2,000
  per day      2,000 * 86,400               =  172,800,000
  per year     172,800,000 * 365            = 63.07 billion

RAW
  per day      172,800,000 * 16 B           = 2.765 GB
  per year     2.765 * 365                  = 1,009 GB  (~1.0 TB)

LOSSLESS AT 1.37 B/POINT
  per day      172,800,000 * 1.37 B         = 236.7 MB
  per year     236.7 MB * 365               = 86.4 GB
  saving                                    = 91.4 %

A terabyte a year becoming 86 GB is the difference between a storage line item somebody notices and one nobody does. It is also, at typical object-storage prices, a change of a few hundred dollars a year at most — which is the more important observation. Raw storage is rarely the expensive part of telemetry. Query cost and ingest request cost usually dominate, and the ingest cost derivation shows where the money actually goes.

The reason to compress anyway is that the compressed representation fits in memory. Gorilla’s entire argument was that a 26-hour window of production monitoring data compressed to a size that fits in RAM, and a query against RAM is orders of magnitude faster than one against disk. Compression here is a latency technique that happens to save disk.

Lossy: downsampling and dead-banding

Lossless encoding cannot beat the information content of the stream. If you need a further order of magnitude, you must decide what to throw away, and the two standard mechanisms discard different things.

Temporal downsampling

Keep full resolution for a recent window and progressively coarser aggregates for older data. Continuing the derivation above with a three-tier policy:

POLICY
  0-7 days      full 10 Hz
  8-90 days     1 Hz  (mean, min, max, count per second = 4 values)
  91-365 days   1/60 Hz (same 4 aggregates per minute)

TIER 1  7 days   * 236.7 MB/day               = 1.66 GB
TIER 2  83 days  * 236.7 MB * (4/10)          = 7.86 GB
TIER 3  275 days * 236.7 MB * (4/600)         = 0.43 GB
                                                --------
        one year retained                       9.95 GB
        against 86.4 GB flat                    88 % less

Keeping min and max alongside the mean is what makes this survivable. A mean alone erases exactly the transients that a fault investigation needs, and the person querying six-month-old data is almost always looking for a spike.

Dead-banding and swinging door

Report-by-exception transmits a value only when it differs from the last transmitted value by more than a set band. It is a compression scheme applied at the source, so it saves radio energy as well as storage, which is why it is ubiquitous in industrial telemetry. The swinging door algorithm generalises it: it keeps a point only when a straight line from the last kept point can no longer represent every intervening sample within a stated tolerance, so a slow linear ramp costs two points regardless of length.

The cost is that the reconstructed series has an error bound rather than being exact, and downstream statistics computed on unevenly spaced points are wrong unless they are time-weighted. A mean over dead-banded samples over-weights periods of rapid change, because those periods contributed more points.

Choosing what to keep

The decision is not really about compression ratios, it is about which questions you intend to answer later. Three rules that hold up:

  • Never lossily compress before the model that needs the detail. A vibration model looking for a bearing signature at 1.2 kHz cannot work from a one-second mean, and the loss is permanent. Compression policy has to be set per channel and per purpose, not per system.
  • Keep raw around the events. A practical compromise is aggressive downsampling everywhere with full-rate retention for a window either side of every alarm. It costs almost nothing because alarms are rare, and it preserves the data anyone will ever ask for.
  • Sampling rate beats compression. Halving the sampling rate halves the data before any encoder runs, and if the signal has no content above the new Nyquist limit it costs nothing at all. That decision, and how to know whether it costs you anything, is the sampling rate trade-off.