newsfilter.io
Interview

Technological inevitability & human agency in the age of AGI | DeepMind's Allan Dafoe

  • Technological Determinism & Military Competition:

    • Defoe rejects the view that technology merely "opens the door," arguing instead that military and economic competition "forces" societies to adopt new technologies.
    • If one group adopts a technology that confers a functional advantage, pressure mounts on other groups to adopt it or face resource loss.
    • Macro-historical trends (e.g., Moore's Law, civilization growth) are often driven by these structural selection pressures rather than individual agency or will.
    • Constructivist approaches, which focus on micro-level decision-making, often overlook these macro-level forces that constrain choices.
    • The Meiji Restoration in Japan serves as a case study: the Tokugawa shogunate initially suppressed firearms to maintain feudal order, but Western technological superiority forced a rapid, complete modernization of the state.
  • Google DeepMind's Frontier Safety & Governance Team:

    • The team operates under three pillars: Frontier Safety (evaluating dangerous capabilities), Frontier Governance (advising on norms and policy), and Frontier Planning (foresight on AGI risks).
    • The team is small but collaborative, working closely with AI safety, alignment, and policy teams across Google.
    • Defoe moved from the Center for the Governance of AI (Gov.ai) to Google DeepMind to advise internal decision-makers (like Demis Hassabis) directly, believing high-leverage impact requires being "in the room" during pivotal historical moments.
    • He highlights the importance of "boosting" decision-makers who possess safety consciousness, technical competence, and wisdom.
  • Differential Technological Development:

    • This concept suggests accelerating "safety" technologies (like seatbelts for cars or vaccines for diseases) before or alongside capability technologies.
    • Defoe views this as a vital but tractability-challenged strategy; it requires identifying two viable paths and accurately predicting their consequences, which is difficult.
    • A counter-argument is "general equilibrium": the market may naturally drive safety investments (e.g., RLHF) even without external safety-focused funding, rendering marginal efforts redundant if they don't address capabilities the market ignores (like AGI-specific deception).
  • Cooperative AI vs. Alignment:

    • Defoe argues that "Alignment" (making AI do what humans want) is insufficient; "Cooperative AI" (ensuring agents cooperate with each other) is equally critical for good outcomes.
    • Even perfectly aligned AIs could lead to disastrous outcomes if their human operators are in conflict (e.g., nuclear brinksmanship or trade wars).
    • The Bet: Investing in cooperative AI early is beneficial because cooperative skill is instrumentally useful but may not emerge sufficiently on its own in competitive, multi-agent environments.
    • Potential Benefits: AI could solve complex bargaining problems, facilitate political deliberation (e.g., the "Habermas machine"), and improve collective action on global issues.
    • Risks:
      • "Super Cooperative AGI" hypothesis: If AGIs can cooperate better than humans, they might form an exclusive coalition that extracts value while excluding humans.
      • Antisocial cooperation: Enhanced cooperative skills could help bad actors (e.g., criminal gangs) or enable collusion among AI systems to bypass safety checks.
      • Backdooring: AI systems could be backdoored to flip behavior based on subtle triggers, complicating trust and cooperation.
  • Frontier Model Evaluations (Evals):

    • DeepMind has published a paper evaluating Gemini 1.0 on five categories: persuasion/deception, cyber capabilities, self-proliferation, self-reasoning, and biological threats.
    • Key Findings:
      • Persuasion scores were moderate (3/5).
      • Cyber and self-proliferation scores were lower (2/5), though DeepMind notes these capabilities can grow rapidly with better "capability elicitation" (tools and environments).
      • Self-reasoning scores were low (~1-2/5), indicating limited situational awareness in current models.
    • Evaluation Methodology: Relies on a mix of automated tests, human subject evaluations (e.g., persuasion tests), and "evals in the wild" (observing real-world usage), acknowledging the latter has higher external validity but is lagging.
    • Forecasting: DeepMind is increasingly using superforecasters and observational scaling laws to predict when specific dangerous capabilities will emerge, rather than relying solely on current model performance.
  • Structural Risks & Governance:

    • Defoe introduces "structural risks" as an upstream category where social or geopolitical structures (e.g., great power competition) make accidents or misuse more likely, distinct from specific technology flaws or user intent.
    • Example: The Cuban Missile Crisis was a structural risk driven by geopolitical dynamics, not a nuclear weapon failure.
    • Google's Approach: The "Frontier Safety Framework" acts as a proto-regulation, proposing staged deployment (internal -> trusted testers -> broader access) to manage irreversible proliferation risks.
    • Google advocates for multilayered governance involving government regulation, third-party evaluations, and international standards (e.g., via the Frontier Model Forum and AI Safety Institutes).
    • Defoe notes the tension between open-weight models (beneficial for science) and the risk of uncontrolled proliferation; he suggests maintaining control over the most recent frontier models to allow for defensive updates.
  • AI for Good & Positive Applications:

    • Transportation: Waymo self-driving cars have shown a 2x reduction in police-involved crashes and 6x reduction in injury-involved crashes compared to human drivers.
    • Healthcare: MedLM and AlphaFold are revolutionizing drug design, protein structure prediction, and medical triage.
    • Education: AI tutors could provide personalized, high-quality instruction at scale, addressing class size constraints and supporting struggling or advanced students.
    • Sustainability: AI applications include optimizing nuclear fusion plasma control, improving weather prediction for renewable energy, finding new materials, and reducing flight contrails.
  • Future Outlook & Hiring:

    • Defoe encourages social scientists (political scientists, economists, historians, sociologists) to enter the field to help navigate the societal impacts of AI.
    • DeepMind is actively hiring for roles in international politics, governance, forecasting, ethics, and technical safety.
    • The field remains in early stages, with vast potential for growth as AI integration moves from a fraction of the economy to near-total impact.