newsfilter.io
Fireside Chat, Interview

Rich Sutton and Khurram Javed: Why AI Models Stop Learning, and How to Start It Again

  • The outlook predicts a paradigm shift where AI evolves from static large language models with frozen weights to systems capable of true continual learning, driven by the expectation that all learning is inherently continuous because agents "always act and we learn."
  • Future systems are expected to bypass finite human-curated data and internet-based training by learning directly from their own experiences in the physical world, leveraging the "big world hypothesis" that the environment is infinitely large to handle complexities like motor friction and wear.
  • Material progress is anticipated in the "next couple of years" with the deployment of right algorithms that support separate step sizes for every weight and the generation of new units, enabling "massively superior continual deep learning" that prevents catastrophic forgetting without requiring frozen models.
  • While current LLMs are viewed as a temporary phase that will "get a good run," the projection for "five to 10 years" includes the ability to run trillion-parameter models on "20 watts" of energy, a capability relying on two orders of magnitude of improvement from Moore's Law.
  • Strategic plans involve Oak Lab starting with a small, highly aligned team of "maybe a handful or two or three" people to develop an architecture for genuine continual learning and abstraction before scaling, contrasting with the industry's current "waste" and energy-intensive approaches.
  • The speaker anticipates that true intelligence will require systems that can "learn new things without catastrophic forgetfulness" by training from scratch rather than applying updates to existing frozen models, addressing the current limitation where agents cannot learn after deployment.
  • A key prediction is that while no single system can learn "infinitely many things," the future will consist of "multiple systems that are learning from their own experience," eventually handling tasks that synthetic data and sim-to-real gaps cannot solve.
  • Risks and barriers include the field being "stuck in a local minnow" due to an unwillingness to pursue new algorithms where performance may not improve immediately, as well as the challenge of maintaining coherence without drifting while "keeping training itself."
  • The long-term trajectory suggests that once machines can plan and reason using self-discovered abstractions, humans may become "irrelevant" regarding task execution, though this is framed as making the world "more exciting and interesting for humans."