newsfilter.io
Interview, Podcast

Ishan Misra: Self-Supervised Deep Learning in Computer Vision | Lex Fridman Podcast #206

  • Self-Supervised Learning Trajectory: Self-supervised learning is anticipated to be a foundational component for future intelligence, emerging with concepts like object permanence, counting, symmetry, and rotation without explicit training, though its exact mechanisms remain unknown or "surprising."
  • Data Efficiency and Generalization Challenges: Current deep learning models struggle to generalize from single or few samples compared to humans, face "catastrophic forgetting" when learning new concepts, and rely on data augmentation that encodes human biases rather than integrated, realistic processes.
  • Autonomous Driving Timeline and Requirements: Solving fully autonomous driving in the US is estimated to take at least five to ten years, with potential extensions to ten or thirty years if full societal integration and AGI-level problem solving regarding human behavior are required.
  • Future of Data and Simulation: The field will shift from curated datasets like ImageNet to uncurated, internet-scale data (e.g., SEER), while simulation is viewed as a difficult but potentially viable path for high-payout scenarios like autonomous driving edge cases.
  • Supervision and Human Interaction Limitations: Human supervision is considered insufficient for large-scale solutions due to inconsistency and inability to label edge cases; consequently, machines must derive supervision from natural signals, requiring embodiment and physical interaction to fully understand the world.
  • Evolution of Learning Methods: Non-contrastive methods (e.g., clustering, self-distillation) are predicted to gain utility over contrastive learning as scaling negative samples becomes difficult, while multimodal learning and "energy-based models" offer frameworks for correlating signals like audio and video to infer common concepts.
  • Infrastructure and Standardization: Architectures like RegNet are optimized for memory and FLOPs efficiency to enable large model training on single GPUs, while initiatives like VSL aim to standardize benchmarks and experiments across the community.
  • Active Learning and Data Engines: Techniques such as active learning and "data engines" (e.g., Tesla's) are expected to become critical for identifying and retraining on edge cases collected from the wild to improve model robustness and reduce labeling costs.
  • Communication and Consciousness Barriers: While AI may display elements of consciousness, emotion, or wit to sustain interaction, self-supervised learning will likely hit boundaries regarding communication interfaces and human knowledge unless "flavor" and mood are integrated to make interactions rich and natural.
  • Theoretical Uncertainties: The field must eventually address "nebulous correctness" where algorithmic guarantees are harder to characterize than in traditional computer science, and the future of intelligence may involve solving complex problems of human society integration rather than just computer vision.