Corpus analysis, 31 full-text readings
What the literature actually says
A corpus analysis of 766 studies, full-text readings of 31 papers, and seven institutional policy documents. This page summarises where the disagreement comes from and how solid each claim is.
The origin of the disagreement
The most explanatory piece of work published in 2026 is not a model but a set of interviews: 25 leading researchers from Google DeepMind, OpenAI, Anthropic, Meta, UC Berkeley, Princeton and Stanford, interviewed in August–September 2025.
- 20/25
- call automating AI research one of the most severe and urgent risks
- 17/25
- expect advanced systems to be held internally, unseen by the public
- Split
- on whether regulatory red lines are a good idea
- Nearly all
- favour transparency-based mitigations
Four mechanisms, not one disagreement
- 01ProximityLab researchers experience the curve first-hand. One participant: “A large part is just having first-person experience of how fast things have gone… inside the labs there's more of a sense of ‘this is what it was like two years ago, and here are all the arguments people brought up at that time, and they just turned out to be false.’”
- 02IncentivesLabs answer to investors, not reviewers; academia answers to reviewers, not investors. One is incentivised to over-promise, the other to over-criticise. As one academic put it: “Skepticism is like a cultural thing in academia… if you're not criticizing, it means that you don't know enough.”
- 03Professional riskThe cost of voicing the same idea is opposite in the two settings. Inside labs, leadership encourages the discussion. In academia, raising intelligence explosion risks being “derided for being a bit of a crackpot.”
- 04Selection effect“Deep believers” migrate from academia into the labs. Part of the gap between the two groups is a consequence of who went where, not of how the evidence was weighed.
The study also finds a clear geographic pattern: Chinese and European research communities are markedly more sceptical. A professor from China reports that “when I talk about intelligence explosion on Weibo, even in China, people don't think it's possible,” and describes Chinese companies as “very application oriented… they just want to use AI to make some money.”
The answer to “when is the singularity?” correlates strongly with which building the answerer sits in. That does not make the answers wrong — but it does show why aggregating timeline forecasts as if they were independent expert judgements is misleading.
The most solid theoretical result
Toby Ord's The Dynamics of Intelligence Explosions (arXiv:2608.14426) is the most important technical contribution of 2026 and it changes the frame of the debate.
- 01Singular growth is much harder than assumedEconomics-inspired models assumed super-exponential growth implies a vertical asymptote. Ord shows this is false. The threshold for super-exponential growth is that Ȧ grows superlinearly in A. The threshold for singular growth is markedly higher: Ȧ must grow faster than all functions of the form A·log(A)·log(log(A))…
- 02The neglected parameter is generation timeThe time to go once around the feedback loop. Singular growth is more about reducing the generation time than about increasing what each loop contributes. With a fixed generation time, no rate of improvement produces a vertical asymptote.
- 03There is a neglected class: super-exponential without an asymptoteMonthly growth rising 10% → 20% → 30% is not the signature of singular growth. In Ord's example, if that pattern continued, A would grow only like e^(0.05·t²) — super-exponential, but never blowing up in finite time.
- 04So local measurement can neither confirm nor rule out an explosionThe blow-up condition is a global property of the relationship between Ȧ and A, not a local property like elasticity > 1. Worse: these systems are exquisitely sensitive near their borderlines, while every empirical measurement of intelligence we can make is rough and noisy.
Ord's concrete proposal: require frontier labs to report their current generation times, especially for pre-training and RLVR post-training. It is the only indicator that is both measurable and relatively closed to manipulation.
The empirical front: is compute the bottleneck?
Whitfill (MIT) and Wu (Yale) built a 2014–2024 panel dataset for OpenAI, DeepMind, Anthropic and DeepSeek, and fitted two CES production functions to estimate the elasticity of substitution (σ) between research compute and cognitive labour.
| Specification | σ | Reading |
|---|---|---|
| Baseline model | 2.583 | Substitutes — a software-only explosion is possible |
| “Frontier experiments” model | −0.103 | Complements — compute becomes the bottleneck |
Two plausible specifications return opposite signs. This is the most honest admission of uncertainty in the field and should not be quoted as a single number. As concrete RSI evidence the paper cites DeepMind's AlphaEvolve, which found algorithmic improvements that cut LLM training time by 1% — real, but modest.
The generalisation front
A cross-generation analysis of 82 approaches across three benchmark versions and the ARC Prize 2024–2025 competitions, as of February 2026.
ARC-AGI: the best system collapses as the version hardens
Best AI score · February 2026 · human performance ~100% on every version
The critical point is not the size of the drop but its consistency: program synthesis, neuro-symbolic and purely neural approaches all show the same 2–3× degradation. That points to a fundamental limit in compositional generalisation, not to one architecture’s shortcoming. Source: Vahdati et al., arXiv:2603.13372 — analysis of 82 approaches.
The critical point is not the size of the drop but its consistency: program synthesis, neuro-symbolic and purely neural approaches all show the same 2–3× degradation from ARC-AGI-1 to ARC-AGI-2. That points to a fundamental limit in compositional generalisation, not to the shortcoming of a particular architecture.
- ARC Prize 2025 winners needed hundreds of thousands of synthetic examples to reach 24% on ARC-AGI-2 — reasoning remains knowledge-bound.
- Cost per task fell 390× in a year (o3 at $4,500/task to GPT-5.2 at $12/task), but the authors note this largely reflects reduced test-time parallelism rather than pure efficiency gains.
- Humans maintain near-perfect accuracy across all three versions.
What the field talks about — term density
| term | 2019 | 2020 | 2021 | 2022 | 2023 | 2024 | 2025 | 2026 |
|---|---|---|---|---|---|---|---|---|
| agent / agentic | 38% | 13% | 24% | 26% | 13% | 17% | 26% | 53% |
| benchmark / evaluation | 15% | 17% | 12% | 26% | 25% | 35% | 42% | 52% |
| recursive self-improvement | 0% | 3% | 0% | 11% | 3% | 6% | 14% | 31% |
| governance / policy | 15% | 10% | 0% | 4% | 17% | 10% | 22% | 24% |
| alignment | 8% | 17% | 0% | 19% | 27% | 19% | 24% | 19% |
| LLM | 0% | 0% | 0% | 22% | 47% | 56% | 40% | 40% |
| economics / growth | 15% | 27% | 18% | 15% | 16% | 13% | 22% | 17% |
| scaling laws | 0% | 3% | 6% | 22% | 8% | 22% | 13% | 8% |
| existential risk | 0% | 13% | 6% | 15% | 5% | 4% | 10% | 4% |
| “singularity” / “intelligence explosion” | 31% | 13% | 24% | 4% | 1% | 3% | 5% | 4% |
- 31%
The field renamed the mechanism
“Singularity” fell from 31% of papers in 2019 to 4% in 2026. Over the same period “recursive self-improvement” rose from 0% to 31%. The concept did not die — its label changed. A speculative-sounding term was traded for an engineering-sounding one.
- 53%
Agents took over the subject
In 2026 more than half the corpus mentions agents. The intelligence-explosion debate no longer runs through an abstract “superintelligence” but through concrete agent architectures.
- 52%
Measurement is the fastest-growing activity
Benchmarking and evaluation went from 25% in 2023 to 52% in 2026. The field moved from “when?” to “how do we measure it?”
- 8%
Scaling laws peaked in 2024 and receded
From 22% to 8%. This is the literature's echo of Sutskever's claim that the age of scaling is over. Meanwhile “world model” barely appears in this corpus at all — LeCun's programme runs on a separate track.
- 160
The discourse itself became an object of study
cs.CY is the third-largest category with 160 papers, examining how singularity claims shape investment and regulation rather than whether they are true.
Contradiction map — who says what, and why
| Dispute | Side A | Side B | Settleable? |
|---|---|---|---|
| Timeline | Altman: threshold crossed · Amodei: 2026–27 | Sutskever: 5–20 years · LeCun: never this way | No. Position tracks institutional position. |
| Mechanism | Scaling plus agents is enough | New research / world models required | Empirically settleable, not yet settled. |
| Software-only explosion | σ = 2.58 → substitutes, possible | σ = −0.10 → complements, bottleneck | Same paper, two specifications. Open. |
| Is an explosion detectable? | The METR curve is accelerating | Ord: local measurement neither confirms nor refutes | Ord's argument is mathematical. Side B is stronger. |
| Is generalisation solved? | 93% on ARC-AGI-1 | 68.8% on ARC-AGI-2, 13% on ARC-AGI-3; 2–3× drop in every paradigm | No. The evidence sits with side B. |
| Should it be banned? | FLI: 72,621 signatures, moratorium | Labs: graduated threshold frameworks | A value judgement; not empirically settleable. |
How solid is each claim?
Charter clause 6.2 requires every numerical claim on this site to carry a robustness grade. Here is that grading applied to the research itself.
Strong — multi-source, measured
- 01Task time horizons are rising exponentially and accelerated in 2024–2025 (7 months → 4 months). METR, multiple releases, open methodology.
- 02Scaling is physically possible through 2030 (~2e29 FLOP). Epoch AI, four constraints modelled separately; electricity binds first.
- 03Compositional generalisation is unsolved. ARC-AGI-2/3, 82 approaches, consistent degradation across all paradigms.
- 04arXiv volume grew 8× in four years and is still accelerating. Direct count.
- 05Frontier labs have accepted the intelligence explosion as a formal risk category. Written into DeepMind FSF v3.1.
Moderate — single source or methodological caveat
- 0120 of 25 researchers call AI R&D automation among the most urgent risks. Single study, non-random sample, single coder — limits stated by the authors themselves.
- 0217 of 25 expect advanced systems to be held internally. Same caveats; but if true, it is the most serious finding for public oversight.
- 03Metaculus: July 2033 for full general AI. Crowd forecast from scored forecasters — but an archive snapshot; the live value may have moved.
Weak — contradictory or framing-sensitive
- 01Whether a software-only intelligence explosion is possible. Same data, two estimates with opposite signs.
- 02Expert survey medians (2047). Framing produces a 69-year swing. Cannot be used as a point estimate.
- 03Company executives' timelines. Simultaneously forecast, investor communication and regulatory lobbying.
Wrong, or badly posed
- 01“We are in the singularity.” Under Vinge's definition (progress exceeding human comprehension) this is not a measurable claim; under Altman's own definition (“wonders become routine”) it is close to a tautology.
- 02“Local measurements show an explosion.” Contradicted by Ord's mathematical result: the blow-up condition is a global property.
- 03“The METR curve is evidence of an intelligence explosion.” Ord: task horizons can go to infinity without other measures of intelligence following.
What this research cannot do
- The corpus is query-shaped. Built from 17 thematic queries; it is not a neutral sample of the field. Term shares are within-corpus.
- 31 papers were read in full; 766 were not. For the remaining 735 studies only title, authors and abstract were used.
- No citation-network analysis. Papers were ranked by topical relevance, not by citation data.
- Company statements come from secondary sources. Executives' 2026 verbal remarks were taken from news reporting; primary audio or video was not verified.
- Metaculus values are archive snapshots (10 and 23 February 2026); live values may have changed.
- No Turkish or regional literature was surveyed. The corpus is limited to English-language arXiv.
- Ord's model is a preprint, not peer-reviewed; its argument is mathematical and can be checked on its own terms.