newsfilter.io
Interview

Jitendra Malik: Computer Vision | Lex Fridman Podcast #110

  • Solving complex computer vision problems is predicted to require "some time" due to the cerebral cortex's processing complexity, with video recognition expected to lag roughly "10 years" behind 2009-level static object recognition despite anticipated progress "over the next few years."
  • Fully autonomous driving is considered unlikely "in the near future" due to the inability of current systems to handle sophisticated cognitive reasoning required for approximately "0.01 percent" of cases, while medical diagnosis systems will likely need to provide specific "error bounds" and decision quality metrics.
  • Achieving long-term knowledge accumulation and human-level intelligence is not expected "in the next 20 years" due to "unknown unknowns" in natural language, which is identified as the "hardest nut to crack," though "rapid progress" may still surprise the field.
  • Future systems will require "new methods of learning" evolving beyond current supervised techniques, potentially mimicking child development through "active learning" in simulation or robotics to reduce the massive data requirements of current computer vision models.
  • By "2020," raw computing capacity will match the human brain, yet systems will remain significantly "power hungry" compared to biological ones, necessitating "life realistic" simulation environments that accurately model forces, masses, and haptic interactions.
  • Progress in long-form video understanding will depend on implementing learning versions of "1970s" style schemas involving goals and intentionality, moving away from the current "Turing test" framework toward benchmarks comprising a "list of 10 different tasks."
  • While future AI systems may become "fundamentally black boxes" for interpretability, this is deemed acceptable given current high-performance models, though risks of harmful biases in deployed settings like hiring and self-driving remain a "continuous problem."
  • Algorithms in recommender systems on platforms like YouTube and Facebook are already described as "super intelligent" with control over billions of people, and future breakthroughs may rely on mimicking human learning rather than manual hand-coding.