AGI Timeline Predictions and Their Track Record
5 min read · updated August 3, 2026
The field has been predicting human-level machine intelligence for seventy years, which sounds like a large evidence base and is not. Most of those predictions cannot be scored, and the reasons they cannot are the useful part.
What makes a prediction scoreable
A forecast that can be graded needs four things, and most published AGI predictions have one or two of them.
- A resolution criterion. A stated, third-party checkable condition. “Human-level intelligence” is not one; “a system passes this specified evaluation under these conditions” is. The seven candidate definitions in what would count as AGI resolve at wildly different times, so a prediction without one is not a claim about the world yet.
- A date or a distribution. A point estimate is scoreable but crude; a distribution over years is much more informative and much rarer, because it forces the forecaster to state how uncertain they are.
- A stated confidence. “By 2040” and “ninety percent chance by 2040” are graded differently. A forecaster who never attaches confidence cannot be shown to be calibrated or miscalibrated.
- A public timestamp. Otherwise the prediction can be revised silently, and a revised prediction with the old date attached is not evidence of anything.
Apply those four to almost any famous AGI prediction and it fails at the first. That is not a rhetorical point against the forecasters — many were writing informally, in interviews, and would not claim otherwise. It is a reason that the aggregate track record of the field is much weaker evidence than it appears, in both directions: it cannot establish that predictions are systematically too optimistic any more than it can establish the reverse.
The historical record, and what it shows
What can be said, from documents anyone can read, is that the early decades contain several confident predictions of near-term human-level machine intelligence from central figures in the field, and that those predictions did not come true on the stated horizon. The 1955 Dartmouth proposal suggested that significant progress on several core problems — language, abstraction, self-improvement — could be made by a small group over a summer. Herbert Simon predicted in 1965 that within twenty years machines would be capable of doing any work a person can do. Marvin Minsky made comparably confident statements in the same period. The 1973 Lighthill report to the UK Science Research Council reached a sharply negative assessment of the field’s progress against its claims, and funding contracted afterwards in what is usually called the first AI winter.
Three things follow, and it is worth being precise about which is which. Empirically: those specific predictions were wrong on those specific horizons. Structurally: they were mostly unscoreable in the sense above, so what the record demonstrates is a pattern of over-confident informal statements rather than a measured bias with an effect size. And as a caution about inference: the failure of 1960s predictions is weak evidence about 2020s ones, because the predictions were made about a different technical situation with different evidence available — which is precisely the outside-view-versus-inside- view problem discussed in why forecasting AI progress is so hard.
There is a second pattern in the record that cuts the other way and is quoted less often. Several capabilities were predicted to be decades away or permanently out of reach and arrived earlier than the sceptics expected — competitive play in Go, fluent open-domain translation, protein structure prediction at useful accuracy. A record containing both over-optimism about the destination and under-estimation of particular waypoints does not support a simple correction in either direction.
Where current timelines come from
| Source | Description |
|---|---|
| expert surveys | Researchers are asked when they expect defined milestones. Broad sampling, and highly sensitive to question wording — the same respondents give substantially different answers to differently framed versions of the same question, which the survey authors themselves report. |
| forecasting platforms | Metaculus and prediction markets aggregate many forecasters with public resolution criteria and track records. The criteria are the strength here. The weakness is that long-horizon questions attract few forecasters and, in real-money markets, the cost of capital makes distant questions poorly priced. |
| bio-anchors | Estimate the compute needed for transformative capability by analogy to biological reference points — a brain's operations, or the compute of evolution — then ask when that compute becomes affordable. Explicit about its assumptions and highly sensitive to them; the analogy itself is the contested step. |
| trend extrapolation | Fit a curve to a capability measure and read off when it crosses a threshold. Only as good as the measure, and capability measures saturate, get contaminated, and get replaced. |
| insider statements | Public statements from people at frontier labs. They have information nobody else has, and a direct commercial interest in the answer. Both facts are relevant and neither cancels the other. |
Reading an expert survey properly
Expert surveys of AI researchers are the most cited source of timeline numbers and the most misread. Three things about their methodology change what the number means, and all three are reported in the papers themselves rather than being external criticisms.
Framing effects are large. The published surveys include a deliberate manipulation: some respondents are asked when machines will be able to perform all tasks better than humans, others when all human occupations will be fully automatable. These are close to logically equivalent, and the aggregate answers differ substantially — which means the elicited number is partly a fact about the wording.
Aggregation hides disagreement. A median across respondents whose individual answers span a century is a summary of a field that does not agree, and reporting it as the field’s expectation misrepresents that. The spread is the finding.
Expertise in building is not expertise in forecasting. Training a model well does not confer calibration about macro trajectories, and unlike the forecasting platforms, survey respondents have no track record on this class of question. This is a limitation the survey authors state; it is often dropped when the number is repeated.
None of that makes the surveys worthless. They are the best available evidence about what the people closest to the work believe, and thechanges between successive waves are more informative than any single figure, because the wording is held constant across waves. Read them for the direction of movement and the width of the distribution, not for the median.
Four checks on any timeline claim
- Which definition? If it is unstated, the claim has no content yet. Ask before evaluating.
- What would falsify it, and by when? A forecast whose proponent would not change their mind under any near-term observation is a disposition, not a prediction.
- What does the forecaster gain if it is believed? Not a refutation — incentive is not evidence of error — but it does tell you which direction to check hardest, and it applies symmetrically to people selling near-term capability and to people selling its impossibility.
- Is the reasoning visible? A derivation you can attack is worth more than a number you cannot. Bio-anchor style models are frequently criticised, and their virtue is that they can be: the assumptions are written down and each one can be swapped.
This page carries no date of its own, deliberately. Any specific timeline stated here would be stale before it was read, and would add one more unscoreable prediction to a literature that already has plenty.