newsfilter.io

MIT 6.S094: Deep Learning for Human Sensing

  • The course will apply deep learning to human understanding in driving contexts, specifically focusing on computer vision for pedestrians and cyclists, with a human-centered approach prioritized over full autonomy (L4) which is predicted to be more than two decades away.
  • Data collection is identified as the most critical success factor, with the speaker expecting large-scale distributed compute and storage to process over 5 billion images, noting the MIT dataset contains 400,000 miles while Tesla data includes a billion miles.
  • Algorithm development plans involve using 3D convolutional neural networks and RNNs/LSTMs to capture visual characteristics and temporal dynamics, targeting tasks like glance classification, body pose estimation, and semantic segmentation.
  • Future systems are expected to rely on large-scale data rather than specific algorithms, utilizing calibration-free methods to ensure robustness across vehicles, with a prediction that fully autonomous vehicles removing the steering wheel are far in the future.
  • Safety risks are highlighted by 2014 statistics showing 3,000 distraction-related deaths and 400,000 injuries, with 31% of fatalities involving drunk drivers, 23% of nighttime drivers testing positive for medication, and nearly 3% involving drowsiness.
  • Human factors include the finding that humans are highly capable drivers who may over-trust or misuse technology, such as bypassing sensors, necessitating systems that perceive driver mental states.
  • Data collection methodologies at MIT involve 25 vehicles (21 with Tesla Autopilot) equipped with cameras on the driver for face video, fish-eye cameras for body pose, and scene cameras, capturing over 1,000 miles daily and 10 hours at specific intersections.
  • Specific technical goals include annotating high-lighting variation and occlusion frames to achieve low 90% accuracy in glance classification, and using eye dynamics to detect cognitive load with 86% accuracy on real-world data.
  • Driver monitoring plans include real-time classification of gaze regions (road, mirrors, clusters) and emotion recognition for broad categories or specific interactions, noting 42 facial muscles are involved in expressions.
  • Upcoming initiatives include the release of a human-centered autonomous vehicle in Boston streets in March 2018, a course on deep learning for humans at CHI 2018, and a "deep traffic" competition requiring 65 mph speeds by a specified deadline with a high performer award for 70 mph.
  • Global business and AI robotics classes are planned for the spring, alongside a deep learning introductory course, with industry contributions acknowledged from NVIDIA, Google, Amazon, Toyota, and others.