Skip to content

Choosing a Sampling Rate for IoT Data Collection

10 min read · updated August 11, 2026

Sampling too fast wastes energy and storage, and both are recoverable later. Sampling too slowly destroys information permanently and does it silently, producing a signal that looks clean and is wrong. The two errors are not symmetric, and that asymmetry should drive the decision.

The floor: Nyquist and aliasing

The sampling theorem, stated by Claude Shannon in “Communication in the Presence of Noise”, published in the Proceedings of the IRE in January 1949, says that a signal containing no frequency components at or above B is completely determined by samples taken at a rate strictly greater than 2B. Half the sampling rate is the Nyquist frequency; content above it is not merely lost.

It is worse than lost, because it reappears somewhere else. A component at frequency f sampled at rate fs appears in the sampled data at |f - k*fs| for whichever integer k brings it into the range below fs/2. The classic worked case:

50 Hz mains interference, sampled at fs = 60 Hz
  Nyquist limit = 30 Hz, and 50 > 30, so it aliases
  k = 1:  |50 - 60| = 10 Hz

The recorded data now contains a clean, strong 10 Hz oscillation
that does not exist in the physical signal. Nothing downstream can
distinguish it from a real 10 Hz component.

This is the failure mode that makes under-sampling different in kind from every other data quality problem. There is no residual, no gap, no flag. The series looks perfectly plausible, a model trained on it learns the artefact, and the only way to discover it is to know the physics of the signal or to resample the same phenomenon faster and compare.

So the floor is set by the highest frequency actually present in the physical signal, not by the highest frequency you care about. Human gait has essentially all its energy below 20 Hz, so 50 Hz sampling is comfortable. Rolling-element bearing defect frequencies are typically hundreds of hertz to a few kilohertz, and their harmonics run higher, so vibration monitoring for bearing faults commonly needs several kilohertz — two orders of magnitude more than the temperature sensor on the same machine.

The theorem also has a consequence that surprises people coming from the other direction: sampling far above the Nyquist rate buys almost nothing in fidelity. Once the rate exceeds twice the highest present frequency, the samples already determine the continuous signal completely, and further increases add data without adding information. What oversampling does buy is easier filtering — the anti-alias filter has more room to roll off — and a lower noise floor after averaging, since averaging M samples reduces uncorrelated noise by the square root of M. Sampling at 1 kHz and averaging groups of ten to produce a 100 Hz series is a legitimate and common design; sampling at 1 kHz and storing all of it because more felt safer is not.

The anti-alias filter is not optional

Because content above Nyquist folds down rather than disappearing, it must be removed before the analogue-to-digital converter, in analogue hardware. A digital low-pass filter applied to already-sampled data cannot help: the folding happened at the sampling instant and the aliased energy is now indistinguishable from real low-frequency signal.

The same rule applies when you reduce a rate in software. Decimating a 1 kHz series to 100 Hz by keeping every tenth sample aliases everything between 50 and 500 Hz down into the retained band. Correct decimation applies a digital low-pass filter at the new Nyquist limit first and then discards samples — which is precisely why SciPy exposes scipy.signal.decimate as a distinct function rather than expecting you to slice the array. Down-sampling a stored series for a retention policy has the same requirement, and taking a mean over each interval is a crude low-pass filter that at least does not simply alias.

Volume and current at three rates

One device, three candidate rates. Every input below is a stated assumption; the arithmetic is what transfers.

ASSUMPTIONS
  sensor            3-axis accelerometer, 16-bit per axis
  payload           6 bytes per sample (timestamps implicit at fixed rate)
  acquisition       0.5 ms awake per sample at 6 mA
  sleep current     5 uA
  battery           225 mAh coin cell, usable capacity assumed 100%

DAILY VOLUME  = rate * 6 B * 86,400 s
  10 Hz    10 * 6 * 86,400  =   5,184,000 B  =  5.18 MB/day
  50 Hz    50 * 6 * 86,400  =  25,920,000 B  = 25.92 MB/day
 200 Hz   200 * 6 * 86,400  = 103,680,000 B  = 103.7 MB/day

DUTY CYCLE    = rate * 0.0005 s
  10 Hz    0.005    50 Hz  0.025    200 Hz  0.10

AVERAGE CURRENT = 6 mA * duty + 5 uA * (1 - duty)
  10 Hz    6000*0.005 + 5*0.995  =  30.0 +  4.98  =  35.0 uA
  50 Hz    6000*0.025 + 5*0.975  = 150.0 +  4.88  = 154.9 uA
 200 Hz    6000*0.100 + 5*0.900  = 600.0 +  4.50  = 604.5 uA

LIFETIME = 225,000 uAh / average current
  10 Hz   225,000 / 35.0   = 6,429 h = 268 days
  50 Hz   225,000 / 154.9  = 1,453 h =  61 days
 200 Hz   225,000 / 604.5  =   372 h =  15.5 days

Note the shape of the result. Volume scales exactly linearly with rate, as it must. Current does not quite, because the sleep term is constant — but the sleep term is only 5 µA, so once the duty cycle is above about half a percent the relationship is effectively linear too, and lifetime is inversely proportional to rate. Going from 10 Hz to 200 Hz turns nine months into two weeks.

The other thing the table shows is that at 10 Hz the sleep current is 14 percent of the total. Below that rate, optimising the acquisition further stops helping, because you are paying for the device merely existing. That is the point at which the design question changes from “how often do we sample” to “how deeply can we sleep”.

The radio changes the answer

The derivation above deliberately excludes transmission, because including it changes the conclusion completely. A Bluetooth Low Energy radio drawing on the order of 10 mA while active dwarfs a 6 mA ADC acquisition, and it has to be awake for far longer per byte than the ADC is per sample.

The consequence is that at 200 Hz, streaming 103.7 MB a day off a coin cell is not merely expensive, it is impossible — the radio would need to be on essentially continuously and the battery would last hours. This is what forces the architecture: sample fast locally because the physics demands it, then reduce on the device to something transmissible. A 200 Hz stream reduced to one feature vector per second is a bandwidth reduction of two orders of magnitude, and it is the reason the preprocessing split usually lands on the device side for high-rate sensors and the cloud side for slow ones.

Choosing the rate

  • Start from the signal, not the budget. Establish the highest frequency with real content — from the physics, from the manufacturer, or by recording a burst at the highest rate the hardware supports and looking at the spectrum. Everything else is negotiation below that ceiling.
  • Take a margin over 2B. The theorem says strictly greater than twice; practice uses two and a half to five times, because real anti-alias filters have a finite roll-off and do not reach zero exactly at the cutoff.
  • Burst rather than sample continuously. For rotating machinery, a 2-second burst at 5 kHz every ten minutes captures the spectral content and costs a duty cycle of 0.3 percent. Rate and continuity are separate decisions and conflating them is the most common way a power budget gets blown.
  • Sample the fast channels fast and the slow ones slowly. A single system-wide rate set by the most demanding sensor is extremely common and wastes most of its energy on channels whose content is below 0.1 Hz.
  • Record the rate as data, not as configuration. If a firmware update changes the rate and the pipeline does not know, every frequency-domain feature computed downstream silently shifts. That is a schema drift problem and it is a common one.