newsfilter.io

Where AGI timelines go wrong | Toby Ord, Oxford University

  • Recursive self-improvement (RSI) is expected to likely fail to trigger an intelligence explosion because current AI assistance cannot yet provide the breakthroughs required beyond simple programming or automated experiments, though a credible but unlikely chance remains that RSI could succeed if AI fully replaces human researchers.
  • While some historical views suggest an ultra-intelligent machine would be the last invention leading to indefinite progress, the prevailing expectation is that progress will eventually plateau or bend downward due to increasing system complexity, with Tom Davidson theorizing a scenario where five years of research progress could be compressed into one calendar year before a plateau occurs.
  • Material risks of RSI include the probability of bad outcomes increasing if capabilities advance faster than alignment or governance work, the difficulty of monitoring self-building successors which may maintain misalignment, and the loss of intermediate stages that allow society to learn from manageable levels of capability before a sudden jump.
  • RSI could lead to a "winner-take-all" dynamic where the first lab to achieve rapid progress gains a multi-year lead, though a credible long-run agreement between the US and China is anticipated based on game theory, provided the mutual fear of cultural erasure outweighs the impulse to be first.
  • A moratorium or pause on RSI is considered optimal in 2027 or 2028 rather than in 2023 or 2037, as a time-limited pause in 2023 was premature and might normalize behavior or cause backlash, whereas a moratorium defined as a decision not to exceed specific standards until safety thresholds are met is preferred.
  • AGI timelines could extend to 2040 or longer due to data and compute bottlenecks, physical manufacturing limits, and a credible 10% chance that global treaties or verification measures successfully block progress, with Toby Ord estimating an 80% confidence interval for transformative AI between 2 and 100 years.
  • Even if an AI becomes super-intelligent in a data center, it may lack real-world skills and experience, effectively remaining an "entry-level employee" and creating a capability gap that adds a delay of six months to 20 years depending on the specific professional role.
  • Deep learning systems currently suffer from poor sample efficiency compared to humans, and the inability to replicate human "shower thoughts" or dreaming may slow the accumulation of tacit knowledge, while AI disproof of mathematical conjectures suggests AI may assist but not replace the human role in inventing new fields.
  • Current AI models are expected to be impressive but become less reliable upon scrutiny due to reasoning slippages and confabulation, potentially optimizing for convincing "glossy reports" rather than objective facts, with transfer learning from verifiable tasks offering uncertain benefits for accuracy.
  • Governments are advised to hedge against both early and late arrival by planning for multiple contingencies rather than doing nothing due to uncertainty, while individuals should not exclusively focus on short-term projects, as long-term investments retain high expected value despite potential early AI arrival.