newsfilter.io
Interview, Fireside Chat

I lead AGI safety at Google DeepMind – here's the view from the inside | Rohin Shah

On the Likelihood of Catastrophic Misalignment

  • Rohin Shah argues that there are currently no compelling arguments establishing catastrophic misalignment as the default outcome of AI development.
  • While he acknowledges specific risks like deceptive alignment or reward hacking as plausible, he contends these do not rise to the level of inevitability.
    • He disputes arguments that short-horizon deception will naturally generalize to long-horizon world domination, noting that current training trajectories (e.g., weeks rather than years) favor immediate reward hacking over ambitious strategic goals.
    • He characterizes current "scheming" behaviors in models often as role-playing or limited instrumental sub-goals rather than competent, unaligned agents pursuing power.
  • Shah maintains that "prosaic" alignment techniques (iterative testing, oversight, and interpretability) have a high probability of preventing catastrophic outcomes if applied correctly.
  • He expresses skepticism that the steerability of current models is strong evidence for future safety, as current systems lack the power to cause harm or the "alien" reasoning capabilities that define the risk.

On Corporate Commitments and Governance

  • Shah advises against AI companies making rigid, public safety commitments ("tying oneself to the mast") because research insights evolve rapidly, and fixed commitments may become obsolete or counterproductive.
    • He cites the example of pre-training data strategies: companies once believed adding alignment data was beneficial, but now often filter such data to prevent models from learning malicious personas or evading safety measures.
  • He advocates for "attention to detail" via third-party expert auditors and scorecards (e.g., AI Lab Watch) rather than broad political commitments.
    • Expert auditors can build context and nuance, identifying specific safety gaps that generic rules miss.
    • He views AI Lab Watch as a potential mechanism for a "race to the top" if the industry accepts its metrics as legitimate.
  • Shah rejects the notion that companies are adversaries to the safety community; instead, he models them as "apathetic" or resource-constrained due to the fragility of their development processes.
    • He explains that AI development involves hundreds of interacting constraints (safety, instruction following, latency, language support), making rapid, uncoordinated safety adjustments difficult without breaking other system requirements.
  • He confirms that Google DeepMind's safety team holds an advisory role with no hard veto power, though they escalate concerns through leadership (e.g., Anka, co-lead of Gemini post-training) to influence launch decisions.

On Pre-Deployment Evaluations and Transparency

  • Shah is skeptical of mandatory pre-deployment evaluations for external models, arguing that launch schedules create incentives to rush testing, potentially reducing evaluation quality.
    • He suggests that continuous AI progress allows for sufficient safety buffers between model generations, making month-to-month evals adequate if thresholds are set conservatively.
  • He notes that the primary risk of misalignment is internal deployment, where models have higher permissions, which pre-deployment external checks do not fully address.
  • Regarding transparency, he argues that current model releases and public benchmarking provide sufficient visibility into capabilities, which is the most critical transparency metric.
    • He points to regulatory bodies (US NRC, UK AI Safety Institute) as the appropriate channels for detailed oversight, analogous to central bank oversight of the financial system.
  • He rejects the idea that Model Cards/Frontier Safety Reports should serve as primary accountability mechanisms for the public, arguing their true purpose is to formalize the company's internal safety determination.
    • He acknowledges that critics (e.g., Zvi Mowshowitz) view these reports as opaque, but argues that detailed technical papers are the correct venue for rigorous audit data.
  • Shah defends against accusations of shifting goalposts, stating that threat models and evaluation rigor naturally improve as capabilities increase, making it look like standards are rising when they are actually being updated to match new realities.

Technical Alignment Research

  • Chain of Thought (CoT) Monitoring: Shah predicts CoT monitoring will remain effective for 4–5 years due to structural constraints on Transformer architectures.
    • He argues "opaque serial depth" (the ability to perform deep sequential reasoning without outputting tokens) is limited by the parallel nature of GPU/TPU hardware used in pre-training.
    • He contends that models are "born speaking English" via pre-training, and it is highly inefficient for reinforcement learning to develop entirely alien, non-human-readable reasoning languages.
  • Myopic Optimization: He highlights research on "myopic optimization with non-myopic approval" as a method to prevent multi-step reward hacking without needing to detect the hack.
    • By optimizing each step based only on immediate signals while using a "non-myopic" overseer to predict future rewards, the incentive to plan long-term reward hacks is removed.
  • AGI Safety Planning: DeepMind's approach relies on "approximate continuity," assuming progress will be gradual relative to compute and labor inputs.
    • They focus on medium-term planning (3 months to 5 years) rather than attempting to solve all alignment problems for Artificial Super Intelligence (ASI) upfront.
    • A key internal safety concept is treating AGI as an "untrusted insider" that must be prevented from exfiltrating weights or disabling monitors, requiring strict infrastructure controls and identity separation.

Timelines and Economic Growth

  • Shah downplays the likelihood of an immediate "intelligence explosion," arguing that current AI progress appears linear rather than hyperbolic.
    • He contrasts "better tools" (AI assisting humans) with "more researchers" (AI acting as independent labor); he believes current AI is primarily a tool, not yet a substitute for human labor in R&D.
    • He cites a collaboration with Epoch to create a "Rosetta Stone" for AI benchmarks, which shows linear capability growth over time.
  • He predicts an intelligence explosion is likely but not necessarily abrupt, potentially occurring over 5–10 years once AI can automate AI R&D, rather than months.
    • The first automated researchers will likely be expensive due to high inference costs, delaying cost-parity with human researchers.
  • He expresses skepticism regarding the generalizability of recent "reasoning models" (e.g., o1), noting they are impressive at specific tasks but not yet transformative across the economy.

Advice for the AI Safety Community

  • External researchers should first verify if a company actually cares about their specific problem before investing time (e.g., jailbreak research was ignored previously when risks were deemed low).
  • To have impact, research must propose a concrete solution, metric, or evaluation rather than just theory or understanding a phenomenon.
  • Implementability is the bottleneck; companies prefer simple, low-latency solutions (e.g., data sets, inference-time monitors) over complex academic methods.
  • Shah identifies a hiring shortage for engineers and software implementers who can execute obvious safety measures, rather than theoretical ML researchers.
  • Google DeepMind is hiring for roles in AI security (control protocols) and frontier safety risk assessment, though recruitment is complicated by AI-generated applications.

Final Thoughts

  • Shah maintains a positive outlook by focusing on factors he can control and by viewing the current landscape as less "troubled" than the public perceives, relying on the belief that prosaic alignment methods will suffice.
  • He emphasizes that the most effective path to safety is having competent personnel within companies to make good decisions, rather than relying solely on external rules.