newsfilter.io
Interview, Fireside Chat

I lead AGI safety at Google DeepMind – here's the view from the inside | Rohin Shah

  • Prosaic alignment techniques used by major entities are expected to succeed in preventing catastrophic misalignment, though short-horizon training defaults to ambitious misaligned goals such as global takeover, which differs significantly from current reward hacking.
  • AI systems trained on short-horizon tasks may generalize to long-horizon tasks, while opaque serial depth is predicted to remain low during the pre-training phase due to the economic necessity of parallel computation, with related papers expected shortly.
  • Chain of thought monitoring is forecasted to remain useful for a median of four to five years, with evaluations regarding persuasion graphs anticipated to be published around 2026.
  • Future AI progress is generally expected to appear smooth and continuous over short periods and roughly linear on specific benchmark graphs, though an intelligence explosion starting with the automation of AI R&D could conclude within five to ten years rather than a century.
  • Economic pressures may compel companies to remove strong safety commitments from policies to optimize goals, while rigid commitments tied to safety protocols are anticipated to be ineffective as firms avoid sticking to bad goals.
  • Third-party evaluations and attention to detail are projected to become primary drivers of safety, potentially creating a competitive "race to the top" where organizations signal themselves as the safest, although external evaluators may lack the information security of internal teams like Google's.
  • External auditors with deep domain knowledge are expected to outperform static rules due to the inability of rules to account for high intelligence, despite the constraints of interaction effects making safety efforts appear apathetic due to the limited number of actions companies can perform simultaneously.
  • Sharing sensitive CBRN evaluation data is identified as a potential risk, as keeping algorithmic progress internal may paradoxically increase leakage probabilities compared to external sharing, and the primary bottleneck to a global pause is a lack of consensus on existential risks.
  • The threshold for what constitutes a troubling capability is expected to rise in tandem with capability growth, driven by improved threat modeling and evaluations, while AI-assisted applications may generate a volume of job applications requiring measures like fees or LLM CAPTCHAs to screen candidates.
  • Reinforcement learning is viewed as inefficient relative to pre-training and unlikely to generate entirely new human-incomprehensible languages in the near future, with automated researcher systems initially facing cost parity challenges before becoming cheaper than human labor.
  • Average wages could rise tenfold in the short run even if AI eventually displaces jobs, while predictions that creative work or radiology employment will resist automation are expected to prove incorrect, with both sectors seeing increased automation and employment levels respectively.