Skip to content

What It Costs to Ingest a Million Sensor Readings a Day

10 min read · updated August 11, 2026

A million readings a day sounds substantial and is, in storage terms, almost nothing. What it can cost a great deal of is messages, and whether it does comes down to a batching decision made in device firmware.

Four cost terms, priced differently

Telemetry platforms bill on axes that scale with different things, and conflating them is why estimates come out wrong by two orders of magnitude.

  • Messages. Priced per message or per fixed-size metered block regardless of content. Scales with how often devices publish, not with how much data they send.
  • Bytes transferred. Priced per GB, and on cellular links priced per MB at rates hundreds of times higher.
  • Storage. Priced per GB-month, and cumulative: it grows with retention, so month twelve pays for everything since month one.
  • Query and processing. Priced per GB scanned, per compute-hour, or per provisioned unit. Scales with how often you read, which is unrelated to how much you write.

Message cost, and why batching dominates

The inputs are stated assumptions. Managed IoT platform message pricing varies widely by vendor, region and metering rule, so replace the price with your own; the ratio is the durable result.

ASSUMPTIONS
  readings                 1,000,000 per day
  reading, compact binary  40 bytes
  reading, JSON with tags  180 bytes
  message price            $1.00 per million messages  (assumed)
  metering                 per message, no size banding (assumed)

UNBATCHED: one reading per message
  messages/day    1,000,000
  cost/day        1,000,000 / 1e6 * $1.00      = $1.00
  cost/month      * 30                          = $30.00
  cost/year                                     = $365.00

BATCHED: 100 readings per message
  messages/day    10,000
  cost/day        10,000 / 1e6 * $1.00          = $0.01
  cost/month                                    = $0.30
  cost/year                                     = $3.65

The bytes are identical. The bill differs by 100x.

Thirty dollars a month is not, on its own, a crisis. The reason this matters is that it scales with fleet growth while the storage term barely does, so an architecture that publishes one message per reading becomes the dominant line item at a hundred times the volume, and by then the fix is a firmware change across a deployed fleet.

Batching is not free. It adds latency equal to the batch interval, so a system with an alerting requirement has to either keep the batch window under the alert deadline or publish urgent readings immediately outside the batch — the standard pattern, and worth designing in from the start. It also increases the loss window: a device that fails mid-interval loses everything buffered. And some vendors meter in fixed-size blocks, so a large batch counts as several messages, which caps the saving at the block size and should be checked before assuming a hundred-fold gain.

The $1.00 per million figure is an illustrative assumption, not a quote. Published IoT messaging prices differ substantially between providers and several meter by kilobyte blocks rather than per message, which changes the arithmetic above materially. Check the current pricing page for the platform you are actually using.

Storage cost, derived over a year

ASSUMPTIONS
  readings           1,000,000 per day
  stored size        40 bytes per reading, uncompressed columnar
  object storage     $0.023 per GB-month
                     (AWS S3 Standard, first 50 TB tier, us-east-1,
                      published price at the time of writing)
  retention          12 months, nothing deleted

VOLUME
  per day     1,000,000 * 40 B          = 40 MB
  per month   40 MB * 30                = 1.2 GB
  per year    1.2 * 12                  = 14.4 GB

COST, cumulative because storage accrues
  month n stores n * 1.2 GB
  year total = 0.023 * 1.2 * (1+2+...+12)
             = 0.023 * 1.2 * 78
             = $2.15 for the first year

WITH TIME-SERIES COMPRESSION at ~1.4 bytes/point
  per year    1,000,000 * 365 * 1.4 B   = 511 MB
  cost                                   under $0.20 for the year

Two dollars a year. That is the number worth internalising, because it reframes the whole exercise: raw storage of sensor data at this scale is free in any practical sense, and optimising it is a misallocation of effort. Compression is worth doing for query speed and for memory residency, as sensor stream compression sets out, and essentially never for the storage bill.

The per-GB-month figure comes from Amazon’s published S3 pricing page, which is the thing to re-check rather than this page: object storage pricing has moved repeatedly, and managed time-series database pricing is a different and much higher order of magnitude, frequently billed per provisioned unit rather than per GB. If you are storing telemetry in a managed TSDB rather than in object storage, expect the storage term to stop being negligible — that is a choice made for query performance, and it should be costed as such rather than as storage.

One caveat that does bite at this scale: many object stores charge per request as well as per byte, and a pipeline writing one small object per reading generates a million PUT requests a day. At typical per-thousand-request pricing that dwarfs the two dollars of storage entirely, and it is the same batching lesson as the message term. Writing hourly Parquet files instead of per-reading objects reduces those million requests to twenty-four, and has the side benefit of producing files a columnar query engine can actually scan efficiently.

The term that actually grows

Query cost is the one that surprises people, because it is unrelated to how much you wrote and entirely a function of how you read. A dashboard refreshing every thirty seconds and scanning a month of data reads far more bytes per day than the devices wrote.

ASSUMPTIONS
  dashboard scans   one month of one metric per refresh
  month of data     1.2 GB uncompressed
  refresh           every 30 s, 8 hours a day
  scan price        $5.00 per TB scanned  (assumed)

refreshes/day   8 * 3600 / 30                    = 960
bytes scanned   960 * 1.2 GB                      = 1,152 GB/day
cost/day        1.152 TB * $5.00                  = $5.76
cost/month                                        = $173

Against $0.30/month of ingest and $2.15/year of storage, the
dashboard is the entire bill.

The fixes are all about not scanning raw data to answer a question that does not need it. Precomputed rollups at the resolutions the dashboard actually displays turn a 1.2 GB scan into a few hundred kilobytes. Partitioning by time and by device means a query for one device reads one partition. Caching a result for the refresh interval removes most of the repeats outright. All three are standard and all three are commonly skipped because the ingest side was the part that got sized.

Sizing your own

The whole estimate is four multiplications, and the value is in keeping the terms separate rather than in any particular price:

messages/month = readings/month / readings_per_batch
message cost   = messages/month * price_per_message

bytes/month    = readings/month * bytes_per_reading
transfer cost  = bytes/month * price_per_byte      # ~0 on wifi, large on cellular

storage cost   = sum over retained months of (cumulative GB * price_per_GB_month)

query cost     = queries/month * bytes_scanned_per_query * price_per_byte_scanned
  • Compute the message term first. It is the one most often forgotten and the one most sensitive to a firmware decision.
  • Check whether transfer is metered. On a cellular fleet it dominates everything else, as the derivation in edge and cloud inference cost shows.
  • Model storage cumulatively. A flat monthly figure understates a growing archive by roughly half over the first year.
  • Estimate reads, not just writes. The number of dashboards and their refresh rates is a cost input, and it is the one that grows with the number of people using the system.
  • Reduce at the source where you can. Halving the sampling rate or applying report-by-exception at the device removes cost from every term at once, which is what makes the sampling rate decision the highest-leverage one available.