Webinar, Interview
What the hell happened with AGI timelines in 2025?
Timeline Shifts: Forecasters for transformative AGI moved from late-2024/early-2025 optimism to second-half 2025 pessimism, with prediction markets extending the arrival of strong AGI from July 2031 to November 2033.
- Industry Sentiment: Early 2025 confidence, exemplified by Sam Altman's declaration of knowing how to build AGI and Demis Hassabis's 3–5 year estimate, gave way to a "swing back" where forecasts blew out further than pre-reasoning model expectations.
- Narrative Drivers: The "AI 2027" scenario (fully automated research and recursive self-improvement) generated massive engagement but was internally viewed as slightly more optimistic than the writers' actual expectations for a superhuman coda.
Technical Factors Reducing Timeline Expectations:
- Failure of Generalization: Reinforcement learning in checkable domains (math, coding) did not generalize effectively to uncheckable, real-world autonomy tasks (e.g., booking flights) as previously hoped.
- Impact: Industry staff updated timelines toward longer horizons after realizing the "free" path to rapid capability gains via generalization was ruled out.
- Anthropic's Approach: New tools like Claude Opus 4.5 and Claude Code rely on specific training for autonomy rather than emergent generalization from reasoning tasks.
- Diminishing Returns on Inference Scaling: Over two-thirds of the performance gain in 2024–2025 reasoning models resulted from increased "thinking time" (inference scaling), a resource with hard physical limits.
- Compute Constraints: The world lacks sufficient chips to sustain exponential growth in thinking time (e.g., moving from 1 minute to 10 or 100 minutes) without prohibitive costs.
- Economic Reality: Costs for advanced agents now approach or exceed human labor rates (e.g., ~$100/hour for complex coding tasks), making further scaling of thinking time economically irrational in the near term.
- Inefficiency of Reinforcement Learning: Scaling reinforcement learning for reasoning yielded only modest returns compared to the massive compute required.
- Compute Penalty: Toby Ord estimates compute efficiency for this method may be one-millionth of the efficiency seen in the "predict next token" pre-training era.
- Feedback Sparsity: Models receive binary right/wrong signals after generating vast numbers of failed attempts, making it difficult to identify exactly where reasoning went astray.
- Failure of Generalization: Reinforcement learning in checkable domains (math, coding) did not generalize effectively to uncheckable, real-world autonomy tasks (e.g., booking flights) as previously hoped.
Autonomy and Real-World Utility:
- Performance Gap: A widening disconnect emerged between demo-level capabilities and actual workplace productivity gains, summarized by Dwarkesh Patel as models improving fast in capability but slow in usefulness.
- Continual Learning Failure: AI models fail to learn incrementally like humans (quickly adapting via few-shot examples), causing them to plateau in capability once deployed rather than improving over time.
- Research Automation Limits: Fully automating AI R&D is hindered by factors beyond software engineering.
- Non-Code Bottlenecks: AI research involves significant non-software activities (experimentation, theory) that code-writing automation cannot resolve.
- Dilution of Gains: As AI writes 90–95% of code, it may only maintain current research speeds rather than accelerating them, a factor that pushed the "AI 2027" creators to extend their own timelines by 1–2 years.
Economic and Adoption Realities:
- Revenue vs. Forecast: Bullish short-timeline forecasts underestimated revenue growth; predicted annual revenue of $16 billion for major AI firms was surpassed by an actual ~five-fold increase to $30 billion.
- Cost-Performance Trends: While frontier models are expensive, costs are falling rapidly (e.g., OpenAI's $4,500 per question on ARC-AGI dropped to $11 in one year), making near-frontier models highly cost-effective.
- Profitability: Companies are profitable on each new paying user and are growing revenue at 5x annually, contradicting narratives of imminent bankruptcy.
- Capability Verification: The EPOC Capabilities Index indicates AI progress has actually doubled since April 2024, refuting claims that progress has stopped or GPT-5 was a failure.
Forward-Looking Statements and Constraints:
- 2028–2032 Critical Window: This period is identified as "make-or-break" for AGI due to resource constraints.
- Compute Saturation: By 2032, the AI industry may consume a massive fraction of global compute and electricity, leaving little "slack" for expansion.
- Financial Risk: Training costs for single models could reach $1 trillion to $10 trillion, creating high-stakes investment risks if ROI from labor replacement isn't proven.
- Revised AGI Forecasts: Long-timeline skeptics (e.g., Gary Marcus, Yann LeCun) have converged on a ~10-year timeline (approx. 2034), a shift from "decades away" to "decades" being a shockingly short timeframe for preparing for societal upheaval.
- Plausible Scenarios:
- 2029–2030: Fully automated AI R&D becomes plausible if current trends continue without major roadblocks.
- Slower Takeoff: A significant risk remains for a much longer, slower grind driven by physical chip manufacturing rates rather than algorithmic breakthroughs.
- 2028–2032 Critical Window: This period is identified as "make-or-break" for AGI due to resource constraints.