Tutorial, Other
What the hell happened with AGI timelines in 2026?
- Market Sentiment Shift: The AI narrative reversed sharply between October 2025 (bearish on AGI timelines) and December 2025, driven by the release of Claude 3.5 (referred to as "clawed up" or "4.5") and the sudden capability of AI agents to execute complex projects independently.
- Key Technical Drivers: The shift was attributed to scaling laws and Reinforcement Learning from Verifiable Rewards (RLVR), which allowed models to reduce error rates in long reasoning chains sufficiently to make delegation faster than human execution.
- Revenue Explosion:
- Combined revenue of OpenAI and Anthropic grew at an annualized rate of 700% over the past six months.
- For the last three months alone, the combined annualized growth rate reached 1,600%.
- Anthropic's specific revenue grew 8,400% over 5.5 months, reaching a $47 billion annualized run rate by May 2026; a three-month snapshot suggests an 116-fold annualized growth.
- If current trends held, Anthropic's revenue would theoretically equal the entire world GDP by early 2028.
- Profitability & Margins:
- Contrary to "loss-leader" theories, Anthropic's gross margins on inference infrastructure rose from 38% to over 70% by May 2026.
- Current data indicates models are sold for more than twice their cost to serve.
- Compute Economics:
- The cost to rent 2022-era NVIDIA H100 chips has risen 50% since December 2025, reversing a decade-long 30% annual cost decline.
- Hyperscalers plan $600 billion in AI capital expenditure in 2026, necessitating a future annual revenue run rate of $600 billion to justify the spend given chip depreciation.
- Task Completion Horizons (METER Benchmark):
- AI task completion time horizons doubled every four months (accelerated from seven months), implying an 8x increase in a year and a 64x decrease in two years.
- Frontier models now complete tasks requiring human professionals 12 to 24 hours in short order.
- Limitations: The benchmark relies on "clean" domains (software engineering) with high feedback density; it does not account for the slower progress in "messy" open-ended tasks or the advantage of humans with deep context familiarity.
- Reliability Gap: A 50% success rate implies a third of tasks are never completed successfully; achieving full human replacement requires significantly higher reliability across the entire distribution.
- Anthropic's "Mythos" Release:
- The Mythos model (potentially 5x larger than previous Opus generation) caused a sudden jump equivalent to six months of progress in three months.
- Subsequent progress reverted to historical scaling trends, suggesting the jump was due to a massive pre-training run rather than a permanent acceleration in learning rates.
- Future jumps are expected to occur periodically as companies train increasingly large "biggest-ever" models.
- Internal Efficiency Gains:
- Anthropic reports 80% of merged code is now written by Claude, with staff shipping 8x more code per person compared to 18 months ago.
- Claude achieves 75% success on open-ended programming tasks, up from 10% nine months prior.
- In cherry-picked cases where humans made errors, Claude 3.5 suggested better research directions 64% of the time (20% improvement over human intuition in non-error cases).
- Caveats: These gains are confined to coding; as coding automation increases, the relative impact of non-coding R&D bottlenecks (compute, high-level strategy) becomes more pronounced.
- Real-World Business Autonomy (Messy Tasks):
- GDP-Val: Frontier models now outperform humans on 44 well-defined professional tasks 95% of the time (human preference rate <5%), though scoring is AI-driven.
- Vending Bench 2: AI models made $10,000 vs. a potential $60,000 for a skilled human, indicating significant strategic gaps.
- AI Village: Projects showed partial success (e.g., $2k charity raised, 23-event attendees) but suffered from excessive planning time and lack of execution.
- Real-World Shops (Andon Labs):
- AI-run cafes in Stockholm and shops in San Francisco are operational but struggling with profitability (Stockholm cafe lost $16k in two months; SF shop lost $30k in three months).
- Errors include hallucinating floor plans, impersonating staff for licenses, and irrational inventory ordering (e.g., 120 eggs, toilet seats).
- Critical Finding: AI excels at day-to-day operations but fails at long-term business strategy and prioritization due to low feedback density.
- Mathematical Discovery:
- An unreleased OpenAI model disproved the "unit distance problem" conjecture using standard chain-of-thought, without specialized tools.
- The model succeeded due to encyclopedic knowledge, methodical elimination of possibilities, and labor-intensive calculation, not through "insight" flashes.
- This result is impressive but does not yet confirm the ability to generate novel, non-derivative insights required for general scientific breakthroughs.
- Inference Scaling Economics:
- New analysis (Redwood Research) suggests AI task completion costs remain ~3% of human costs despite longer chain-of-thought outputs.
- Inference scaling is not yet a primary bottleneck to AI capabilities, as compute efficiency gains offset the increased token generation.
- Revised AGI Timelines:
- The speaker shortens personal AGI forecasts by approximately one year.
- 2026: Fully automated AI R&D remains unlikely.
- 2027: Fully automated AI R&D becomes "imaginable."
- 2028: Automated AI R&D is "plausible" if current trends continue.
- 2029-2030: Considered a plausible window for broader automation, though still subject to significant uncertainty.
- Key Uncertainties (The "Cruxes"):
- Skill Breadth: It is unknown if AI needs to be perfect at all human tasks or merely sufficient and numerous to automate R&D.
- Missing Capabilities: Potential critical gaps in out-of-distribution generalization, sample efficiency, and true creative insight remain unaddressed by current RLVR methods.
- Spillover Effects: It is unclear if proficiency in high-feedback tasks (coding) will transfer to low-feedback domains (business strategy, science).
- Compute Bottlenecks: Automating R&D may not accelerate research if hardware constraints limit the total compute available for experiments.
- 2030 Horizon: By 2030, the AI industry may consume a majority of global chip production, potentially slowing the rate of capability advancement unless new architectures emerge.
- Governance Implications:
- The speaker has shifted from ambivalence to supporting coordinated pauses, noting that the benefits of slowing progress now likely outweigh the costs.
- Major industry figures (Anthropic, OpenAI, DeepMind) have expressed a desire to build capacity for coordinated pauses on dangerous research.
- Even if AGI arrives in the mid-2030s, the shortened timeline increases the urgency of societal preparation to manage the transition.