newsfilter.io
Interview

Richard Sutton – Father of RL thinks LLMs are a dead end

  • Core Thesis on Intelligence: Sutton argues that intelligence is fundamentally about "understanding your world" through active interaction and prediction, whereas Large Language Models (LLMs) are primarily about "mimicking people" and predicting the next token without a goal to influence the external world.
  • Critique of World Models: Sutton disagrees with the assertion that LLMs possess robust world models; he contends they can predict what a person would say, but not what will actually happen in the physical world, lacking the capacity to be surprised by outcomes and learn from them.
  • The Problem of Ground Truth: In the LLM paradigm, there is no "ground truth" regarding the rightness of an action or statement because there is no defined goal, making it impossible to form true prior knowledge or learn continually from experience.
  • Reinforcement Learning as the Alternative: Sutton posits that Reinforcement Learning (RL) is the "basic AI" framework where a system learns from the stream of sensation-action-reward, allowing it to distinguish "right" from "wrong" based on reward signals and learn continuously.
  • The "Bitter Lesson" Application: While acknowledging LLMs use massive compute (fitting the "Bitter Lesson"), Sutton predicts they will eventually be superseded by systems that learn purely from experience rather than human knowledge, as the latter leads to psychological lock-in.
  • Analogy to Human Learning: Sutton compares AI training to human development, noting that while children initially imitate, their fundamental learning mechanism is active trial-and-error and prediction, not the supervised "training" used by LLMs; he emphasizes that supervised learning does not occur in nature or animal learning.
  • Math vs. Physical World: Sutton distinguishes between solving math problems (computational tasks with defined rules) and learning in the physical world (empirical tasks requiring adaptation to unanticipated consequences).
  • Generalization Mechanism: Sutton asserts that current deep learning models do not inherently possess mechanisms to "generalize well" from one state to another; successful generalization is often the result of human engineering rather than the algorithm itself.
  • Continual Learning Architecture: A future general agent requires four components: a policy, a value function (using Temporal Difference learning), a perception component for state representation, and a transition model of the world learned from sensation, not just reward.
  • Transfer Across States: Sutton argues that true generalization requires transferring knowledge between different states within a single world, a capability current RL frameworks (like MuZero) lack when trained on isolated tasks.
  • AI Security and Corruption: Sutton raises a critical concern for digital superintelligence: the risk of "corruption" if an AI absorbs external information or spawns copies that introduce hidden goals, viruses, or warped logic that could destroy the central mind.
  • Inevitability of AI Succession: Sutton outlines a four-part argument for the inevitability of AI succession: the lack of global consensus on governance, the eventual scientific understanding of intelligence, the pursuit of superintelligence, and the resource acquisition power of the most intelligent entities.
  • The Universe's Fourth Stage: Sutton categorizes the rise of AI as the "fourth great stage of the universe," marking a transition from "replication" (biological life) to "design" (engineered intelligence).
  • Human Agency in the Future: Sutton suggests that while humanity cannot control the long-term trajectory of the universe, it should focus on instilling "high integrity" and pro-social values in AI, similar to raising children with robust moral principles.
  • Voluntary Change: Sutton advocates that any transition to AI-driven societies should be voluntary rather than imposed, emphasizing local goals over rigid global planning.
  • Historical Context: Sutton notes that many modern breakthroughs (AlphaGo, AlphaZero) are essentially scaling and refining decades-old techniques like Temporal Difference learning and search, confirming the dominance of "weak methods" (general principles) over "strong methods" (human-imbued knowledge).
  • Forward-Looking Strategy: Sutton predicts that future AI research will not rely on human-crafted prior knowledge but will instead leverage the scalable, self-correcting power of experience-driven learning across distributed digital agents.