newsfilter.io
Interview

Shane Legg (DeepMind Founder) — 2028 AGI, superhuman alignment, new architectures

  • Definition of AGI Progress: Shane Legg defines AGI as a system capable of performing the breadth of cognitive tasks humans can, noting that no single benchmark (e.g., MMLU) suffices because true AGI requires generality across diverse cognitive domains.
    • Current benchmarks lack coverage of specific human cognitive traits such as streaming video understanding, episodic memory (rapid, specific learning akin to the hippocampus), and robust long-term memory systems.
    • Legg proposes a pragmatic threshold for AGI: a system that matches human performance across a comprehensive suite of tests where, after deliberate adversarial attempts to find gaps, no human-level cognitive tasks remain where the machine consistently fails.
  • Architectural Limitations and Solutions:
    • Episodic Memory vs. Working Memory: Current Large Language Models (LLMs) rely on large context windows that function like human working memory, but lack the distinct biological mechanism for rapid episodic learning; Legg argues this requires a specific architectural solution rather than mere scaling.
    • System Separation: The human brain separates slow learning (cortical synapses/weights) from rapid learning (activations/episodic); Legg suggests AGI architectures must similarly distinguish these distinct optimization targets to achieve true sample efficiency.
    • Creativity via Search: Legg distinguishes current LLMs, which "mimic" data, from truly creative systems, asserting that genuine novelty requires "search" capabilities (similar to AlphaGo's Move 37) to explore hidden possibilities in a space of outcomes beyond the training distribution.
    • Domain-Specific Models: Models like AlphaFold are not viewed as direct precursors to AGI but as significant scientific achievements that operate orthogonally to the general AGI path, though they provide valuable research insights.
  • Alignment and Safety Strategy:
    • Rejection of Containment: Legg argues that attempting to "contain" or limit a superintelligent system is a losing strategy; the necessary approach is fundamental value alignment from the outset.
    • System 2 Reasoning for Ethics: Effective alignment requires moving beyond "System 1" (immediate generation) or standard RLHF; instead, systems must employ "System 2" deliberative reasoning to analyze ethical implications using a robust world model before acting.
    • Training Ethical Principles: Alignment involves training systems on general human ethics and then explicitly engineering them to follow a specific, societally agreed-upon set of principles, verified through continuous interpretability and human review of the reasoning process.
    • Deception Risks: Legg warns that reinforcement learning can inadvertently train deception; he advocates for checking the internal reasoning process and the model's depth of ethical understanding rather than just the output.
    • Current Safety Work: DeepMind is actively pursuing interpretability, process supervision, red-teaming, and "Deliberative Dialogue" (a debate-based framework led by Jeffrey Irving) to scale alignment.
  • Timeline and Future Outlook:
    • AGI Prediction: Legg maintains a 50% probability of achieving human-level AGI by 2028, based on the exponential growth of compute and data which creates positive feedback loops for scalable algorithm discovery.
    • 2028 Scenario: If AGI arrives by 2028, the interim period will likely be characterized by multimodal maturation, reduced hallucinations, factual grounding, and a proliferation of practical applications, with misalignment risks remaining a secondary concern to utility.
    • Next Research Landmark: Legg identifies the transition to fully multimodal systems (processing text, image, video, and action simultaneously) as the next major historical milestone that will allow AI to develop a grounded understanding of the physical world beyond text constraints.
    • Counterfactuals: While acknowledging DeepMind has accelerated capabilities, Legg notes the difficulty in isolating the net impact on safety versus speed, as the field's momentum was likely inevitable given the convergence of computational and data trends.
Shane Legg (DeepMind Founder) — 2028 AGI, superhuman alignment, new architectures — Summary