newsfilter.io
Lecture, Conference Presentation, Fireside Chat

MIT AGI: Building machines that see, learn, and think like people (Josh Tenenbaum)

  • Current deep learning technologies are projected to be near-capable of pattern recognition but insufficient for human-level intelligence, which requires modeling the world, common sense, and flexible general-purpose capabilities.
  • The near-term focus (five to ten years) will target visual intelligence by addressing gaps in understanding the human visual system, building on deep network successes to develop architectures integrating bottom-up perception, a cognitive core for space and objects, a "brain OS," and symbolic representations for language and planning.
  • Industry incentives are currently biased toward short-term value propositions over a two-year horizon, though science-based reverse engineering is expected to eventually gain traction as its successes become evident.
  • High-value progress toward advanced AI is predicted to rely on science-based reverse engineering rather than scaling pattern recognition and reinforcement learning, with a specific goal to achieve AGI systems that develop intelligence similarly to humans, starting from a state comparable to a baby.
  • Future systems are expected to utilize "probabilistic programs" combining symbolic representation, probabilistic inference, and neural architectures to enable learning from sparse data, transfer learning, and one-shot learning of generative models.
  • Cognitive models will likely incorporate "game engines in the head" (intuitive physics and psychology) to serve as foundational approximations for common sense, allowing machines to perform inverse planning over physics models to infer goals from observed actions in novel scenes without large training datasets.
  • Fundamental questions regarding consciousness, meaning, language, real learning, culture, and creativity are anticipated to be addressed over a 50-year timeframe, alongside a shift from optimization landscapes to "learning as programming" where internal models are actively refined.
  • Hardware advancements are deemed necessary for brain-scale computation, potentially leveraging video game industry innovations like "Spatial OS" and low-power approximate computing to address energy efficiency bottlenecks for robots and self-driving cars.
  • Individual neurons may function similarly to CPU nodes, implying the human brain contains roughly 10 million cores rather than simply billion simple units, necessitating a convergence of programming languages, Bayesian inference, and machine learning to enable machines to learn programs.
  • A primary risk involves the field settling for merely scaled-up pattern recognition, failing to achieve true general intelligence capable of interacting in the human world, as evidenced by current limitations in complex visual scene understanding.
  • Understanding how emotions are filtered through mental models is considered essential for engineering the full spectrum of human emotion, while software-level cognitive understanding will guide neural circuit researchers in identifying specific computational tasks.
  • Future innovations will require a convergence of academic basic research and industry engineering to build machines capable of living and interacting in the human world, with the hope that industry will eventually take on the engineering side of advanced cognitive systems.