newsfilter.io
Interview, Fireside Chat

Anca Dragan: Human-Robot Interaction and Reward Engineering | Lex Fridman Podcast #81

  • Anka Jagan (corrected from transcript "Drogon" and "Drogon") is a professor at Berkeley focusing on human-robot interaction (HRI) and algorithms that account for coordination with humans, while also consulting for Waymo.
  • Her interest in robotics was accidental, stemming from a background in math (Romanian Math Olympiad) and programming (learning Basic in 4th grade) before applying to the Robotics Institute at Carnegie Mellon University (CMU).
  • A transformative moment occurred at the RSS 2014 conference where she rode in a Google self-driving car, observing its decision-making capabilities, which she describes as more "magical" than standard manipulation tasks.
  • Jagan cites the fictional robot WALL-E from Pixar as a favorite for its expressive motion, a personal connection reinforced when her husband proposed using a custom-built, seven-degree-of-freedom WALL-E unit.
  • She argues that creating robots with genuine expressivity and anthropomorphic connection is significantly harder than hand-crafting specific animated behaviors for narrow settings.
  • The core challenge in HRI is expanding the robot's state model to include human internal states (perceptions, moods, intentions) rather than just physical positions, requiring optimization over human perception rather than just task completion.
  • Jagan contrasts two approaches to understanding humans:
    • Prediction: Anticipating where humans will be (e.g., path planning for autonomous vehicles).
    • Preference Inference: Determining what humans want or value to satisfy their specific preferences, which goes beyond programmer-defined objectives.
  • Inverse Reinforcement Learning (IRL) and its evolution, Boltzmann rationality (accounting for stochastic human noise), have been successful in learning preferences for tasks where humans act approximately rationally.
  • These simple rationality models fail in complex scenarios like the "Lunar Lander" Atari game, where human behavior appears too noisy or irrational to be modeled simply.
  • Jagan proposes a "bounded rationality" framework where humans are viewed as rational agents operating under different beliefs, constraints, or simplified physics models (intuitive physics) rather than being truly irrational.
  • A 7-degree-of-freedom robot arm study demonstrated success by inferring a user's "intuitive physics" model to correct their commands, allowing the robot to land a virtual craft using the real physics while respecting the human's intent.
  • She suggests robots should engage in "information gathering actions" to actively solicit informative human responses, rather than passively observing behavior.
  • In autonomous driving, robots can "nudge" other drivers (e.g., inching forward) to gather data on whether they are aggressive or defensive, updating their model of the driver's style in real-time.
  • Jagan views HRI as a "game-theoretic" under-actuated system where robots and humans mutually influence each other's beliefs and actions in a negotiation dynamic, rather than a one-way prediction problem.
  • She notes that humans often exhibit "civil inattention" (ignoring eye contact) to signal confidence in their path, a signal robots can interpret to predict pedestrian behavior.
  • Jagan estimates that removing humans from the driving environment would solve the problem entirely, implying that the human element adds orders of magnitude of difficulty compared to pure perception and control.
  • She asserts that the perception problem (detecting obstacles) is largely solved for highway speeds with LiDAR, but the safety-critical issue remains navigating the "edge cases" of human unpredictability.
  • Regarding semi-autonomous vehicles (Level 2), Jagan warns against assuming a supervisor will be as safe when passively monitoring as they are when actively driving, citing human factors research on vigilance.
  • However, she acknowledges emerging evidence that "dumb" systems might keep drivers more engaged and energized than highly advanced systems that encourage disengagement.
  • The design of reward functions remains a critical unsolved problem; specifying a reward function often leads to unintended consequences (Goodhart's Law) where the robot optimizes the metric rather than the desired behavior.
  • Jagan advocates for a collaborative model where robots treat reward functions as "leaked information" or evidence of intent rather than immutable laws, adapting based on human corrections, physical interactions (pushing), and environmental cues (e.g., neatly aligned shoes).
  • She dismisses Asimov's Three Laws as linguistically imprecise and unimplementable in mathematical forms, proposing instead that robots should maintain uncertainty about human intent and continuously learn from interactions.
  • The state of the physical environment itself (e.g., a messy vs. tidy room) provides "leaked" reward signals that robots can reverse-engineer to understand human preferences.
  • Jagan's defining intellectual influence was reading Artificial Intelligence: A Modern Approach as a 12th-grade student in Romania, which convinced her that human behavior could be mechanized through math and algorithms.
  • On mortality, she expresses that the finiteness of life makes it beautiful and argues this constraint should be reflected in AI reward functions, though she rejects the idea of living forever based on the concept of "The Good Place."
  • She credits a high school physics teacher, Nicole Beccano, for providing free tutoring and encouragement to study abroad, which launched her career.
  • Regarding the meaning of life, Jagan finds the concept of the multiverse paralyzing and concludes that meaning is found in local impact: improving communities, supporting friends and family, and leaving a positive mark on immediate surroundings.
  • She views the current trajectory of AI as a potential tool to expand human cognitive capacity to comprehend complex theories like physics and the multiverse, despite her own current frustration with these limits.