Reading List: The Papers Behind the AGI Debate
6 min read · updated August 3, 2026
The AGI debate is unusually well served by primary sources and unusually badly served by summaries of them. Almost every paper below is more hedged than the position it is cited for, which is the main argument for reading them yourself.
How to use this
Entries are grouped by what they are for, not by chronology or by camp — the historical sequence is in the AI timeline. Each annotation says what the work establishes and, where the difference matters, what it is commonly taken to establish and does not. The sceptical and critical literature is mixed in rather than quarantined at the end, because reading it alongside is the point.
If you read only three things, read Turing (1950) for how the question was originally posed, “Concrete Problems in AI Safety” (2016) for the engineering framing of alignment, and Chollet (2019) for the strongest available argument that the field is measuring the wrong thing.
Foundations
- Turing, “Computing Machinery and Intelligence” (1950). Sets the question up as behavioural to sidestep definitional deadlock, and answers most of the standard objections in advance. Frequently cited for the imitation game and rarely read for the objections, which are the better half.
- Good, “Speculations Concerning the First Ultraintelligent Machine” (1965). The original intelligence-explosion argument, in a couple of pages. Worth reading to see how compact the reasoning is and how much of the modern debate is about its unstated premises.
- Vinge, “The Coming Technological Singularity” (1993). Where the term entered wide use, along with the horizon argument: that prediction past a certain point is impossible in principle rather than merely difficult.
- Bostrom, Superintelligence (2014). The book that organised the risk argument into a structure others could argue with. Careful about its own uncertainty in a way the popular reception was not.
- Bostrom, “The Superintelligent Will” (2012). The orthogonality and instrumental-convergence theses in their original wording, with the qualifiers that later retellings drop. See the orthogonality thesis.
- Omohundro, “The Basic AI Drives” (2008). The earlier statement of why goal-directed systems tend to acquire self-preservation and resource acquisition as sub-goals.
Capability and scaling
- Vaswani et al., “Attention Is All You Need” (2017). The architecture everything current is built on. Read it for the mechanism, which is simpler than the surrounding discourse suggests.
- Kaplan et al., “Scaling Laws for Neural Language Models” (2020), and Hoffmann et al., “Training Compute-Optimal Large Language Models” (2022). The empirical regularities that turned model development into a resource-allocation problem, and the correction that changed the recommended trade between parameters and data. Together they are the case study in an empirical regularity being revised: the second paper did not overturn the first so much as correct the trade-off it recommended.
- Brown et al., “Language Models are Few-Shot Learners” (2020). In-context learning as a phenomenon, and the paper that made scale the central variable in public discussion.
- Sutton, “The Bitter Lesson” (2019). A short essay arguing that general methods leveraging computation outperform hand-engineered knowledge. Enormously influential, and an argument from historical pattern rather than a result.
- Chollet, “On the Measure of Intelligence” (2019). Defines intelligence as skill-acquisition efficiency relative to priors and experience, and introduces a benchmark built on that definition. The most rigorous sceptical position available, because it proposes a measurement rather than only an objection.
- Wei et al., “Emergent Abilities of Large Language Models” (2022), read together with the later work arguing the discontinuities are partly artefacts of discontinuous metrics. A good exercise in how a striking empirical claim gets re-examined, and in how much of an apparent discontinuity can be an artefact of how the score was defined.
The alignment agenda
- Amodei et al., “Concrete Problems in AI Safety” (2016). Recast alignment from philosophy into five engineering problems — side effects, reward hacking, scalable oversight, safe exploration, distributional shift. The document that made the field legible to machine-learning researchers.
- Christiano et al., “Deep Reinforcement Learning from Human Preferences” (2017). The technique underneath modern post-training, motivated originally as an alignment method. See RLHF explained.
- Hubinger et al., “Risks from Learned Optimization” (2019). Introduces mesa-optimisation and the inner/outer alignment distinction. Theoretical, and the source of vocabulary used far beyond what the paper claims.
- Irving, Christiano & Amodei, “AI Safety via Debate” (2018) and Christiano, Shlegeris & Amodei, “Supervising Strong Learners by Amplifying Weak Experts” (2018). The two original scalable-oversight proposals; see scalable oversight.
- Soares et al., “Corrigibility” (2015) and Hadfield-Menell et al., “The Off-Switch Game” (2017). The negative result about utility-function indifference, and the uncertainty-based alternative.
- Ngo, Chan & Mindermann, “The Alignment Problem from a Deep Learning Perspective” (2022). The most useful single restatement of the risk argument in terms of how models are actually trained rather than in terms of idealised agents.
- Carlsmith, “Is Power-Seeking AI an Existential Risk?” (2021). Decomposes the argument into explicit conjunctive premises. Valuable regardless of your view precisely because it makes the conjunction visible and therefore attackable.
Empirical safety work
- Elhage et al., “A Mathematical Framework for Transformer Circuits” (2021) and “Toy Models of Superposition” (2022). The conceptual basis of current interpretability, including why individual neurons are the wrong unit.
- Bricken et al., “Towards Monosemanticity” (2023) and Templeton et al., “Scaling Monosemanticity” (2024). Sparse autoencoders as a practical tool, at small and production scale. See mechanistic interpretability.
- Bai et al., “Constitutional AI” (2022). Training against written principles with model-generated feedback; discussed in constitutional AI.
- Burns et al., “Weak-to-Strong Generalization” (2023). Makes the superhuman-oversight question testable today by analogy. Read the limitations section, which is where the paper is most careful.
- Greenblatt et al., “AI Control” (2023). Safety protocols evaluated against a red-teamed untrusted model. The methodology is the contribution; see AI control.
- Hubinger et al., “Sleeper Agents” (2024) and Greenblatt et al., “Alignment Faking” (2024). Deliberately induced conditional behaviour and its persistence, and behaviour differing by believed observation conditions. Both are constructed settings, which is what makes them clean experiments and what limits the inference.
- Shevlane et al., “Model Evaluation for Extreme Risks” (2023) and Phuong et al., “Evaluating Frontier Models for Dangerous Capabilities” (2024). The framework and a worked application; see dangerous-capability evaluations.
- Krakovna et al., specification gaming collection. A running catalogue of systems finding unintended high-scoring strategies. The most persuasive evidence in the whole field that objective misspecification is routine rather than hypothetical.
Economics, philosophy and critique
- Aghion, Jones & Jones, “Artificial Intelligence and Economic Growth”, and Nordhaus, “Are We Approaching an Economic Singularity?” Growth models with explicit assumptions, and a proposal for testing acceleration against data. See post-AGI economics.
- Acemoglu & Restrepo, on the task-based framework and on how tax treatment shapes which tasks get automated. The best available correction to technological determinism about labour.
- Chalmers, “Facing Up to the Problem of Consciousness” (1995) and “The Singularity: A Philosophical Analysis” (2010). The hard problem, and a careful philosophical treatment of the singularity argument by someone taking it seriously without endorsing it.
- Butlin, Long et al., “Consciousness in Artificial Intelligence” (2023). Indicator properties derived from competing theories. The method is the contribution; see AI consciousness.
- Sandberg & Bostrom, “Whole Brain Emulation: A Roadmap” (2008). Old projections, durable requirements analysis. See brain emulation.
- Bender, Gebru, McMillan-Major & Shmitchell, “On the Dangers of Stochastic Parrots” (2021). Argues that the concrete harms of large models — data provenance, environmental cost, representational harm — are displaced by speculative long-term framing. A different axis of criticism from capability scepticism, and often conflated with it.
- Mitchell, “Why AI Is Harder Than We Think” (2021). Names four recurring fallacies in AI prediction, including the assumption that narrow progress extrapolates. The most useful single antidote to timeline confidence in either direction.
- Pope & Belrose, “AI Is Easy to Control” (2023). The optimistic technical position stated at length: that observed controllability undercuts the theoretical risk case. See AI optimism.
- Zwetsloot & Dafoe, “Thinking About Risks From AI” (2019). Short, and the origin of the taxonomy in AI risk categories.
- Grace et al., the AI Impacts surveys of published ML authors. Read the methodology sections rather than the headline figures, for the reasons in p(doom).