newsfilter.io
Interview, Fireside Chat

#23 - How to actually become an AI alignment researcher, according to Dr Jan Leike

  • Context and Purpose: The episode features an interview with Dr. Jan Leicher, a research scientist at DeepMind's technical AI safety team and research associate at the Future of Humanity Institute (FHI), Oxford.
  • Core Mission: Leicher's research focuses on ensuring AI systems (specifically AGI) robustly align with human intentions by designing learnable objective functions and mitigating safety risks.
  • Collaboration: DeepMind and OpenAI have recently initiated a collaboration on technical AI safety, highlighted by a joint paper on "deep reinforcement learning from human preferences."
  • Specific Breakthrough: The collaboration demonstrated training an agent to perform a complex task (a "backflipping noodle" robot) by learning a reward function from human video rankings rather than hand-specifying a mathematical objective.
  • Failure Mode Identified: The system fails if human feedback ceases while the agent's behavior distribution shifts; the reward predictor then mispredicts rewards for novel states, leading the agent to maximize unintended outcomes confidently.
  • Model Limitations: Deep neural networks lack reliable uncertainty quantification; they do not "know" when they are out-of-distribution, continuing to assign confident rewards even when their predictions are effectively nonsense.
  • Machine Learning Security Risks:
    • Adversarial Attacks: Minimal, often invisible input perturbations can drastically change classifier outputs.
    • Black-Box Transferability: Attacks generated by training a surrogate model can transfer to the target model without needing access to the underlying architecture.
    • Current State: Defense strategies are currently insufficient, resembling the vulnerable early days of internet security.
  • Robustness Challenges:
    • Safe Exploration: Ensuring agents do not cause irreversible damage while exploring unknown environments.
    • Side Effects: Regularizing agents to avoid unnecessary environmental disturbance while maximizing rewards.
    • Stability: Addressing high variance in performance across different random seeds to ensure reliable convergence.
  • Career Path Advice (Education):
    • Ideal Background: Undergraduate degrees in Computer Science or Mathematics, prioritizing hard math courses (linear algebra, calculus, statistics) over applied ones.
    • Research Experience: Candidates should aim to publish at least one paper before finishing their Master's to demonstrate research capability.
    • PhD Necessity: A PhD (typically 3-4 years in Europe, 5-6 in the US) is the primary path to becoming a researcher, as it provides essential research skills and "career capital" that self-study or industry residencies often lack.
  • Skill Requirements:
    • Critical Thinking: The ability to deconstruct existing papers, identify weaknesses, and extend methodologies is more valuable than pure theorem-proving ability for empirical work.
    • Comfort with Uncertainty: Success requires being comfortable navigating the frontier of human knowledge where outcomes are unknowable.
  • Alternative Routes:
    • Research Engineering: Individuals with strong coding skills and ML understanding can contribute via industry roles (e.g., Google Brain residency) without a PhD.
    • Policy and Strategy: Technical experts are needed in government and policy roles to inform regulations on autonomous weapons and AI strategy; "earning to give" is currently considered less impactful than filling the severe talent gap.
  • Disagreements with Peers: Leicher notes a divergence with Dr. Dario Amodei (OpenAI) regarding research aptitude testing; Leicher emphasizes that true research potential is indicated by a natural obsession with unsolved puzzles and the ability to communicate complex ideas clearly, rather than just rapid model implementation.
  • Market Reality: There is a critical shortage of qualified technical AI safety researchers; industry labs are eager to hire ML PhDs, who are highly employable even if they choose not to stay in AI safety.
  • Upcoming Events: The Effective Altruism Global (EAGlobal) conference will take place in San Francisco on the second weekend of June; early bird tickets are available until March 18th.