Skip to content

Why Forecasting AI Progress Is So Hard

5 min read · updated August 3, 2026

Forecasting has a methodology, it is tested, and it works reasonably well in domains with repeated events and clean feedback. AI progress violates most of its preconditions, and it is worth being precise about which ones.

The reference class problem

The core move in serious forecasting is to find a reference class: a set of similar past events whose outcomes give you a base rate. How long do large infrastructure projects overrun? How often do drugs at phase two reach market? The base rate anchors you before you reason about the specifics, and it is the single most reliable protective measure against over-confidence.

For transformative AI, every candidate reference class is contested and each implies a different answer.

  • Other software. Suggests fast diffusion, near-zero marginal cost and rapid capability gains. Chosen by people who expect fast change.
  • Other general-purpose technologies — electricity, the internal combustion engine, computing itself. Suggests decades of lag between invention and measured productivity effect, because complements have to be built. Chosen by economists.
  • Previous AI cycles. Suggests over-promising followed by contraction. Chosen by long-standing sceptics, and its weakness is that the current cycle differs from earlier ones in a way that matters: the systems are commercially deployed at scale, which was not true in the 1970s or 1980s.
  • Growth-mode transitions — agriculture, industrialisation. Suggests rare, fast, world-changing shifts. Its weakness is a sample size in the low single digits.

There is no neutral procedure for choosing among these, and the choice largely determines the forecast. When two people disagree about AI timelines, they are very often disagreeing about reference class selection while discussing capabilities — which is why the argument rarely converges.

Inside view, outside view, and why both fail here

The outside view — ignore the details, use the base rate — is the standard corrective for planners who are too close to their own project. It fails here for the reason above: there is no agreed base rate to apply.

The inside view — model the mechanism, estimate each step — fails differently. It requires knowing what the remaining steps are. If the scaling hypothesis is right, the remaining steps are largely resource acquisition and the problem is close to a forecasting exercise about supply chains. If capabilities such as continual learning or grounded world modelling are genuinely missing, the remaining steps include unknown research breakthroughs, and the arrival time of an unknown breakthrough has no distribution anyone can estimate. Whether we are in the first world or the second is exactly what the scaling debate is about, so the inside view depends on the answer to the question it is being used to settle.

The thing being forecast keeps changing shape

Forecasting works best when the resolution criterion is fixed and external. This domain fights that in three specific ways.

Benchmarks saturate and are replaced. A capability measure that everything scores near the top of stops discriminating, so the community moves to a harder one. The result is a series of incommensurable measures rather than one long time series, and long time series are what trend extrapolation needs.

Measures get contaminated. When a benchmark is public and models train on scraped text, apparent progress on it partly reflects exposure. This is measurable and is measured — see benchmark contamination — but it means a rising score is not automatically a rising capability.

Scaffolding changes the score without changing the model. The same weights perform very differently with better prompting, tool access, retrieval or sampling strategy. So “what can the system do?” is not a property of the model alone, and comparisons across time are comparisons of moving bundles rather than of models.

The two errors, and who makes each

Both directions of error are well represented, and neither side has a monopoly on rigour.

Over-prediction

The characteristic mechanism is generalising from a demonstration. Something impressive in a controlled setting is treated as capability in the general case, and the gap between the two — reliability, edge cases, adversarial inputs, integration with the systems and processes that would have to change — is where the years go. The related error is extrapolating a hardware curve to a capability conclusion without the intervening argument about algorithms and data.

Under-prediction

The characteristic mechanism is arguing from a current limitation to a permanent one. Systems have repeatedly acquired capabilities that confident arguments held to be out of reach for the approach, and the arguments usually had a common form: identifying something the current method does badly and treating that as a structural property of the method rather than a fact about its current scale or training. The related error is dismissing evidence of capability because the mechanism is unlike human cognition, which is a claim about mechanism answering a question about behaviour.

It is worth noticing that these two errors are not symmetric in their consequences, which is a separate point from their frequency and is often smuggled in as though it settled the empirical question. How costly an error is bears on how to act under uncertainty. It does not bear on how likely it is.

What better forecasting looks like

  • Forecast inputs, not outcomes. Compute available, capital deployed, data accessible and algorithmic efficiency all have actual time series, and they are the terms any capability forecast depends on. Getting them right does not give you the answer, but getting them wrong guarantees the answer is wrong.
  • Forecast capabilities, not the label. “A system that can complete a specified multi-day software task unsupervised” is scoreable. “AGI” is not.
  • Give a distribution and say what would move it. A forecaster who names the observation that would shift them has made a falsifiable commitment; one who does not has expressed an attitude.
  • Track your own record. Calibration is trainable, it is measured by scoring rules such as the Brier score, and it is domain-specific. Someone who has forecast this domain publicly for years and been graded is a different kind of source from someone with expertise in building the systems.
  • Separate the three claims. Capability forecasts, deployment forecasts and impact forecasts have different evidence bases and different lags. A capability arriving is not a product shipping, and a product shipping is not an economy changing — which is the whole subject of AI and economic growth.
Why Forecasting AI Progress Is So Hard · Multigrid