AGI Announcements and Goalpost Movement
5 min read · updated August 3, 2026
Arguments about whether a system is approaching general intelligence are unusual in that both sides can be entirely sincere and still never converge, because the term has no agreed referent and each side is using a version that suits its conclusion.
There is no agreed definition
“Artificial general intelligence” is used, by careful people, to mean at least these different things:
- A system matching competent human performance across most cognitive tasks.
- A system that can learn any new task a human could, from comparable exposure.
- A system able to perform economically valuable work across most occupations.
- A system with autonomous goals and the ability to pursue them over long horizons.
- A system with subjective experience and understanding in the full sense.
These are not refinements of one idea. The third is an economic threshold that could be crossed by a system failing the fifth entirely. The fourth is about autonomy rather than ability. A conversation where one participant means the third and another means the fifth cannot resolve, and neither is being unreasonable.
This is not a failure of rigour by the people using the term. The phrase entered circulation as a way of distinguishing broad ambition from the narrow task-specific systems that dominated the field for decades — it was a direction, not a specification, and it did that job well. The trouble began when a directional term started being used as a threshold, at which point everyone needed it to have a definite value and nobody agreed on which. A word that was useful for saying what a research programme was aiming at is being asked to settle whether the programme has arrived.
The pattern in one direction
The oldest observation in this field is that a capability stops counting as intelligence once it is achieved. It has a name — the AI effect — and the pattern is a matter of public record rather than interpretation.
Chess was the canonical test of machine intelligence for decades; a machine won a match against the world champion in 1997 and chess was promptly reclassified as search. Go was said to require intuition of a kind search could not supply; a program beat a leading professional in 2016 and Go became a pattern-recognition problem. Conversational fluency sufficient to be mistaken for a person was the canonical test proposed in 1950, and it is now met routinely by systems nobody proposes to call generally intelligent.
The reclassification is not merely bad faith, which is why the pattern is worth taking seriously rather than mocking. Each time, we learned something real: that the task did not require what we thought it required. That is genuine knowledge and it is the correct response to discovering a task was more tractable than expected. It also means the definition can never be met, since any met criterion becomes evidence the criterion was wrong.
The structural cause is that we have no way to specify intelligence except by listing things intelligent beings can do, and every such list is a list of tasks. Tasks keep turning out to be solvable by methods that do not resemble what the list was trying to capture. So the criterion falls, the concept survives untouched, and the next criterion gets drawn from whatever is still unsolved — which guarantees the target is always defined as the current frontier.
And in the other
The less-discussed movement runs the opposite way, and it has been just as consequential in the last few years. The bar comes down when a term becomes valuable: capabilities that would previously have been described as a good language model get described as a step toward general intelligence, and the word starts appearing in contexts where “a system that does X well” would have been the accurate description.
The mechanism here is definitional rather than empirical. If “AGI” is stipulated as, say, a threshold on some measurable economic or benchmark quantity, then reaching that threshold is achieving AGI by construction — and the interesting question, whether the system can do the general open-ended thing the phrase originally connoted, has been quietly replaced with an easier one. This is not lying. It is choosing an operational definition, which is normally good practice, applied to a term whose whole meaning was the un-operationalised part.
Who benefits from which movement
Applying this evenly is the only way it is worth applying at all.
| Position | Description |
|---|---|
| raising the bar | Serves anyone whose standing depends on the technology being less impressive than claimed — including researchers whose alternative approach is out of favour, and commentators whose audience wants reassurance. Not only critics: an incumbent can benefit from insisting a rival's system falls short. |
| lowering the bar | Serves anyone raising capital, recruiting, or seeking the regulatory attention that comes with being important. Also serves people who sincerely believe a threshold has been crossed and want the language to reflect it. |
Incentive analysis identifies who has a reason to want a conclusion. It never establishes that the conclusion is wrong, and using it as though it did is the standard failure of this genre — reading both camps critically takes that further.
The argument worth having instead
Every question people want to settle by arguing about AGI can be asked without the term, and each version has evidence attached:
- Can it do this specific task at this reliability, on inputs the builder did not choose?
- Does the capability hold when the surface form is varied without changing the problem?
- Can it acquire a genuinely new skill within a session, or does it need retraining?
- Can it operate over a long horizon without a person correcting the trajectory?
None of these needs a definition of intelligence. Each has an answer that can move. And notably, someone who thinks we are close and someone who thinks we are far will often agree about all four — which is a good sign that the disagreement was never about the capabilities.
One caution before dropping the term entirely. Something in the vicinity of the concept is doing real work in policy and in research planning: whatever would let a system take on open-ended tasks nobody anticipated is a genuine question, and replacing it with a list of benchmarks loses exactly the part that mattered. The recommendation here is not that the concept is empty. It is that within any specific argument, the capability questions are the ones with evidence attached, and a discussion that never reaches them was not about the technology.