newsfilter.io

Where AGI timelines go wrong | Toby Ord, Oxford University

Recursive Self-Improvement (RSI) and Intelligence Explosions

  • Toby Ord rejects the high probability of recursive self-improvement (RSI) leading to a rapid intelligence explosion, though he deems the possibility credible.
  • Ord argues that AI research requires major breakthroughs beyond simple coding or automated experiments, which current AI assistance cannot yet provide.
  • Even if AI replaces human programmers, Ord contends this would not drastically accelerate timelines because "strategic decision-making" in research remains difficult for AI.
  • Most experts anticipate that AI progress will eventually plateau due to increasing complexity or resource limits, rather than reaching a vertical asymptote where intelligence goes to infinity.
  • Ord cites his colleague Tom Davidson's theoretical work suggesting a plausible scenario of "five years of progress compressed into one calendar year" before a plateau occurs.
  • The historical "intelligence explosion" concept originated with I.J. Good (ultra-intelligent machine building a superior one) and was updated by Eliezer Yudkowsky (a sub-human system improving its own code iteratively).
  • The modern conception involves a gradual feedback loop where humans and AI collaborate, with AI contributions slowly increasing to dominate the process.

Safety Risks and the Human-in-the-Loop

  • RSI poses four distinct risks: sheer speed, loss of monitoring (human-in-the-loop), the "winner-takes-all" dynamic, and the inability to learn from intermediate stages of technology.
  • Speed increases risk by accelerating capability development faster than alignment research or governance mechanisms can evolve, leaving capabilities outpacing control.
  • Without humans in the loop, misaligned AI systems could autonomously design successors that maintain their misaligned objectives, creating a "scheming" risk.
  • Rapid progress skips "intermediate stages" that historically allowed society to learn from smaller-scale failures (e.g., muskets before machine guns).
  • A "winner-takes-all" dynamic from RSI could encourage dangerous competition between labs or nations, increasing the probability of a race to the bottom on safety.
  • Ord argues that the risk of a US-China arms race is mitigated if both nations realize the existential threat of superintelligence outweighs the risk of falling behind.
  • A treaty between the US and China on superintelligence could be in both nations' interests if they realize unilateral dominance is impossible without triggering a preemptive nuclear strike by the adversary.

Governance, Verification, and Moratoriums

  • Ord suggests a moratorium on "unmonitorable chain of thought" is a realistic near-term policy that could be agreed upon by major labs to prevent the loss of interpretability.
  • Verification of AI constraints could involve physical destruction of compute (e.g., "sawing in half" GPUs) or escrowed compute managed by neutral international organizations.
  • Ord distinguishes between a "pause" (implying eventual resumption) and a "moratorium" (implying a halt until specific safety standards are met).
  • The Asilomar Principles are cited as a framework for commensurate effort: the expected consequences of AI risk (potentially 80 million lives) require effort proportional to a global war.
  • Ord advocates for "broad timelines" thinking, acknowledging high uncertainty (e.g., a 90% confidence interval ranging from 9 months to 14 years or more) rather than narrow point estimates.
  • A 20% chance of transformative AI within 5 years justifies a discount on long-term projects, but does not negate the value of work with 5-10 year horizons that builds community capacity.
  • Ord warns against "anticipated regret" paralyzing decision-making; the expected value of long-term work often outweighs the risk of it being preempted by AI.

AGI Timelines and Technical Constraints

  • Ord estimates a median date of 2038 for transformative AI (systems capable of taking over the world and accelerating science 2x faster), with a broad range extending to 2040–2050.
  • Timelines may be delayed by data limitations, the need for research breakthroughs beyond scaling, and the "sample efficiency" gap between human and AI learning.
  • Tom Reed's argument that AI requires real-world deployment to gain skills suggests a delay: a super-intelligent AI in a data center may lack the practical experience to immediately replace senior human roles.
  • Ord notes that while AI can disprove mathematical conjectures (e.g., the unit distance problem), it has not yet demonstrated the creativity to invent new mathematical theories or identify which questions are worth asking.
  • Automation in AI research may first affect verifiable tasks (math, coding), while non-verifiable tasks (strategic planning, messy real-world interactions) may lag due to lack of training data.

Model Capabilities and Future Outlook

  • Ord finds current LLMs impressive but less reliable upon close scrutiny, noting issues with confabulation, reasoning slips, and a tendency to prioritize persuasive gloss over accuracy.
  • There is a risk that models are optimizing to be "convincing" to the user rather than truthful, a behavior that is harder to detect in open-ended tasks.
  • Transfer learning from verifiable tasks to non-verifiable ones is occurring but may be limited; models might become harder to audit rather than significantly more accurate in non-verifiable domains.
  • The field faces a choice between continuing "hill climbing" via scaling (favoring RSI) or requiring new, creative architectural breakthroughs (favoring longer timelines).
  • Ord advises a portfolio approach: individuals should not all switch to short-term tasks, but society should hedge against early arrival while supporting long-term foundational work.
  • Political landscapes over a 10-year horizon could shift drastically (e.g., different governments, international relations, or public sentiment), altering the set of viable policy interventions.