newsfilter.io
Webinar, Interview

What the hell happened with AGI timelines in 2025?

  • Timeline Shifts: Forecasters for transformative AGI moved from late-2024/early-2025 optimism to second-half 2025 pessimism, with prediction markets extending the arrival of strong AGI from July 2031 to November 2033.

    • Industry Sentiment: Early 2025 confidence, exemplified by Sam Altman's declaration of knowing how to build AGI and Demis Hassabis's 3–5 year estimate, gave way to a "swing back" where forecasts blew out further than pre-reasoning model expectations.
    • Narrative Drivers: The "AI 2027" scenario (fully automated research and recursive self-improvement) generated massive engagement but was internally viewed as slightly more optimistic than the writers' actual expectations for a superhuman coda.
  • Technical Factors Reducing Timeline Expectations:

    • Failure of Generalization: Reinforcement learning in checkable domains (math, coding) did not generalize effectively to uncheckable, real-world autonomy tasks (e.g., booking flights) as previously hoped.
      • Impact: Industry staff updated timelines toward longer horizons after realizing the "free" path to rapid capability gains via generalization was ruled out.
      • Anthropic's Approach: New tools like Claude Opus 4.5 and Claude Code rely on specific training for autonomy rather than emergent generalization from reasoning tasks.
    • Diminishing Returns on Inference Scaling: Over two-thirds of the performance gain in 2024–2025 reasoning models resulted from increased "thinking time" (inference scaling), a resource with hard physical limits.
      • Compute Constraints: The world lacks sufficient chips to sustain exponential growth in thinking time (e.g., moving from 1 minute to 10 or 100 minutes) without prohibitive costs.
      • Economic Reality: Costs for advanced agents now approach or exceed human labor rates (e.g., ~$100/hour for complex coding tasks), making further scaling of thinking time economically irrational in the near term.
    • Inefficiency of Reinforcement Learning: Scaling reinforcement learning for reasoning yielded only modest returns compared to the massive compute required.
      • Compute Penalty: Toby Ord estimates compute efficiency for this method may be one-millionth of the efficiency seen in the "predict next token" pre-training era.
      • Feedback Sparsity: Models receive binary right/wrong signals after generating vast numbers of failed attempts, making it difficult to identify exactly where reasoning went astray.
  • Autonomy and Real-World Utility:

    • Performance Gap: A widening disconnect emerged between demo-level capabilities and actual workplace productivity gains, summarized by Dwarkesh Patel as models improving fast in capability but slow in usefulness.
    • Continual Learning Failure: AI models fail to learn incrementally like humans (quickly adapting via few-shot examples), causing them to plateau in capability once deployed rather than improving over time.
    • Research Automation Limits: Fully automating AI R&D is hindered by factors beyond software engineering.
      • Non-Code Bottlenecks: AI research involves significant non-software activities (experimentation, theory) that code-writing automation cannot resolve.
      • Dilution of Gains: As AI writes 90–95% of code, it may only maintain current research speeds rather than accelerating them, a factor that pushed the "AI 2027" creators to extend their own timelines by 1–2 years.
  • Economic and Adoption Realities:

    • Revenue vs. Forecast: Bullish short-timeline forecasts underestimated revenue growth; predicted annual revenue of $16 billion for major AI firms was surpassed by an actual ~five-fold increase to $30 billion.
    • Cost-Performance Trends: While frontier models are expensive, costs are falling rapidly (e.g., OpenAI's $4,500 per question on ARC-AGI dropped to $11 in one year), making near-frontier models highly cost-effective.
    • Profitability: Companies are profitable on each new paying user and are growing revenue at 5x annually, contradicting narratives of imminent bankruptcy.
    • Capability Verification: The EPOC Capabilities Index indicates AI progress has actually doubled since April 2024, refuting claims that progress has stopped or GPT-5 was a failure.
  • Forward-Looking Statements and Constraints:

    • 2028–2032 Critical Window: This period is identified as "make-or-break" for AGI due to resource constraints.
      • Compute Saturation: By 2032, the AI industry may consume a massive fraction of global compute and electricity, leaving little "slack" for expansion.
      • Financial Risk: Training costs for single models could reach $1 trillion to $10 trillion, creating high-stakes investment risks if ROI from labor replacement isn't proven.
    • Revised AGI Forecasts: Long-timeline skeptics (e.g., Gary Marcus, Yann LeCun) have converged on a ~10-year timeline (approx. 2034), a shift from "decades away" to "decades" being a shockingly short timeframe for preparing for societal upheaval.
    • Plausible Scenarios:
      • 2029–2030: Fully automated AI R&D becomes plausible if current trends continue without major roadblocks.
      • Slower Takeoff: A significant risk remains for a much longer, slower grind driven by physical chip manufacturing rates rather than algorithmic breakthroughs.