Skip to content

Real-Time Fraud Detection From Event Streams

10 min read · updated August 11, 2026

The model is the easy part. Everything hard about real-time fraud detection is in the feature pipeline: computing an aggregate over the last hour in 20 milliseconds, computing the same aggregate identically during training, and evaluating a system whose ground truth arrives three months after the decision.

This page describes how the modelling and stream infrastructure work. It is not guidance on financial controls, and any production system in this area sits inside a regulatory regime — governing model explainability, adverse-action notices and record retention — that requires review by people qualified in it.

Velocity features are windowed aggregates

The features that carry the most signal in card and account fraud are almost all counts and sums over a recent window keyed by an entity: transactions on this card in the last 10 minutes, distinct merchant countries for this card in the last hour, sum of amounts on this device in the last 24 hours, distinct cards seen from this IP in the last 7 days. Each one is a windowed aggregation keyed on a different entity, and the choice of window shape has consequences that are easy to miss.

A tumbling window is wrong here. If “transactions in the last 10 minutes” is computed as the count in the current 10-minute tumbling bucket, then at 10:00:01 the count resets to zero and a card that made nine transactions between 09:52 and 09:59 scores as clean. That reset is a predictable, exploitable hole in the detector. What you want is a genuinely trailing window, which means either a sliding window with a slide fine enough that the granularity error is acceptable, or an explicit ring of small sub-buckets summed at read time.

The sub-bucket approach is the one most production systems land on because it bounds cost. Keep 60 one-minute counters per key; the last-hour count is their sum, the last-10-minute count is the sum of the newest 10, and one key costs 60 small integers regardless of how many windows you serve from it. A sliding window with a one-second slide over an hour, by contrast, assigns every event to 3,600 windows — the size ÷ slide factor — and that is where streaming feature pipelines quietly become unaffordable.

Distinct counts need different structure again. “Distinct countries in the last hour” cannot be summed from sub-buckets without keeping the sets, so either keep exact small sets — countries are a bounded domain, so this is fine — or a sketch such as HyperLogLog for genuinely high-cardinality domains like distinct card numbers per IP, accepting a few percent error.

Point-in-time correctness

This is the failure that ruins more fraud models than any modelling mistake. Training data is built by joining historical transactions to historical features, and the join must reconstruct the feature value as it was at the moment of the decision, not as it is now.

Get it wrong and you leak the future. Suppose the training set computes “transactions on this card in the last hour” with a query that groups the whole day’s transactions by card. A fraudulent card that was used 40 times in the hour after the transaction you are labelling now shows a high count on that transaction too. The model learns that high velocity predicts fraud with implausible accuracy, validation looks superb, and production performance collapses — because in production the future transactions have not happened yet.

The structural fixes are known and worth stating plainly. Compute training features by replaying the event log through the same aggregation code that serves them online, so there is one implementation and no opportunity for the two to differ. Where a feature store is used, use its as-of join, which takes an entity key and a timestamp and returns the value whose validity interval contains that timestamp. And include the feature computation timestamp in the training row so you can audit that it precedes the event timestamp — a check that takes one line and catches the whole class.

The related trap is availability skew: a feature that is present in training because the batch job had all day to compute it, but arrives too late online and is served as null. The model has never seen null for that feature and behaves arbitrarily. Simulate the online availability deadline when building training data, and treat any feature that misses it as missing during training too.

The labels arrive months later

A transaction declined as fraudulent never gets a label — you prevented the outcome you wanted to observe. A transaction approved and later disputed gets a label when the dispute is raised, which under card scheme rules can be a long way after the event; treat the exact window as something to look up in your acquirer’s current rules rather than a number to assume. Either way, the practical consequence is fixed: today’s label set describes decisions made months ago, under an older model, against an older adversary.

Three things follow. First, any evaluation on recent data is measuring an incomplete label set, and the missing labels are not missing at random — they are disproportionately the frauds not yet discovered, so recent-period precision is systematically overstated. Wait out the maturity period before believing a number.

Second, the training set is censored by the previous model’s decisions. You only observe outcomes for transactions that were approved, so the model is trained on the subset its predecessor let through. The standard mitigation is to approve a small random holdout against policy, which is expensive and is the only way to get unbiased data about the region the current model rejects.

Third, the adversary adapts, so drift is not a background nuisance but the main dynamic. Monitor the distribution of the input features and the score distribution continuously — those are available immediately — rather than waiting for label-based metrics, which are the last thing to tell you anything. The general treatment of that monitoring problem is in drift detection and data quality checks.

Choosing a threshold with a cost matrix

A score is not a decision. The threshold comes from the costs, and writing them down as a matrix makes the choice arithmetic instead of a negotiation. Let cFN be the loss from an approved fraud, cFP the cost of declining a legitimate transaction, and p the model’s estimated probability of fraud. The expected-cost-minimising rule is to decline when p · cFN exceeds (1 − p) · cFP, which rearranges to a threshold of cFP / (cFP + cFN).

The numbers are yours, but the shape of the answer is general: because fraud is rare, the threshold that minimises cost sits far from 0.5, and reporting accuracy on this problem is meaningless — a model that approves everything is accurate to within the fraud rate. Precision at a fixed, operationally tolerable review volume is the metric that corresponds to a real decision, and it is what to report. The definitions are in classification metrics, and the arithmetic of why a rare positive class wrecks precision is worked out with numbers in security event log anomaly detection.

What a stream cannot see

Windowed features describe one entity’s recent behaviour. They are blind to structure across entities: a ring of 400 accounts, each individually unremarkable, all funnelling to two destinations, has no per-entity velocity signal at all. Detecting that is a graph problem — connected components, shared-attribute edges, community detection — and it runs on a different cadence, usually batch, because the structure only becomes visible once enough of it exists. The two approaches are complements rather than alternatives, and a system with only the streaming half will keep missing organised fraud while catching impulsive fraud very well.

The other blind spot is the first event for a new entity. Every velocity feature is zero for a brand-new card, device or account, which is indistinguishable from a well-behaved one. Cold-start entities need features that do not depend on their own history — attributes of the request, similarity to known-bad patterns, reputation of the network the request came from — and it is worth training a separate model for that regime rather than letting the main model interpret a wall of zeros.