newsfilter.io
Interview, Conference Presentation

Some thoughts on the Sutton interview

  • The current reliance on static pre-training using human data is considered inefficient and unsustainable because human knowledge is an inelastic resource, comparable to fossil fuels that served only as a cheap, convenient intermediary before a transition to more advanced paradigms.
  • Present LLMs primarily predict the next token rather than learning true world models or how actions affect the environment, resulting in low sample efficiency where models extract approximately one bit of information per episode despite processing tens of thousands of tokens.
  • Future architectures are expected to enable continual learning and on-the-fly training, eliminating the need for a special, exhaustible training phase and rendering current methods obsolete.
  • While human data has been essential for accumulating knowledge over tens of thousands of years, scaling to massive amounts of human data is predicted to become less helpful, with the first AGI likely bootstrapping itself from scratch without relying on initialization from human data.
  • Imitation learning is viewed as a necessary prior and a short-horizon form of reinforcement learning that aids in building the representations required for true world modeling, similar to how AlphaGo utilized human games before AlphaZero achieved superhuman performance from scratch.
  • Techniques such as test-time fine-tuning, making supervised fine-tuning a tool call, or extending information flow beyond context windows may replicate continual learning and meta-learn the flexibility currently observed only within context limits.
  • The timeline for development suggests LLMs will first achieve Human-Computer Interaction (HCI) capabilities, serving as a foundation for successor systems that will be based on architectures prioritizing continual learning and true world modeling.
  • Critiques regarding the lack of world models and continuous learning are identified as genuine gaps, yet current models undergoing reinforcement learning on ground truth are already developing deep, coherent world representations across domains like biology and history.