Interview, Fireside Chat
#23 - How to actually become an AI alignment researcher, according to Dr Jan Leike
- Context and Purpose: The episode features an interview with Dr. Jan Leicher, a research scientist at DeepMind's technical AI safety team and research associate at the Future of Humanity Institute (FHI), Oxford.
- Core Mission: Leicher's research focuses on ensuring AI systems (specifically AGI) robustly align with human intentions by designing learnable objective functions and mitigating safety risks.
- Collaboration: DeepMind and OpenAI have recently initiated a collaboration on technical AI safety, highlighted by a joint paper on "deep reinforcement learning from human preferences."
- Specific Breakthrough: The collaboration demonstrated training an agent to perform a complex task (a "backflipping noodle" robot) by learning a reward function from human video rankings rather than hand-specifying a mathematical objective.
- Failure Mode Identified: The system fails if human feedback ceases while the agent's behavior distribution shifts; the reward predictor then mispredicts rewards for novel states, leading the agent to maximize unintended outcomes confidently.
- Model Limitations: Deep neural networks lack reliable uncertainty quantification; they do not "know" when they are out-of-distribution, continuing to assign confident rewards even when their predictions are effectively nonsense.
- Machine Learning Security Risks:
- Adversarial Attacks: Minimal, often invisible input perturbations can drastically change classifier outputs.
- Black-Box Transferability: Attacks generated by training a surrogate model can transfer to the target model without needing access to the underlying architecture.
- Current State: Defense strategies are currently insufficient, resembling the vulnerable early days of internet security.
- Robustness Challenges:
- Safe Exploration: Ensuring agents do not cause irreversible damage while exploring unknown environments.
- Side Effects: Regularizing agents to avoid unnecessary environmental disturbance while maximizing rewards.
- Stability: Addressing high variance in performance across different random seeds to ensure reliable convergence.
- Career Path Advice (Education):
- Ideal Background: Undergraduate degrees in Computer Science or Mathematics, prioritizing hard math courses (linear algebra, calculus, statistics) over applied ones.
- Research Experience: Candidates should aim to publish at least one paper before finishing their Master's to demonstrate research capability.
- PhD Necessity: A PhD (typically 3-4 years in Europe, 5-6 in the US) is the primary path to becoming a researcher, as it provides essential research skills and "career capital" that self-study or industry residencies often lack.
- Skill Requirements:
- Critical Thinking: The ability to deconstruct existing papers, identify weaknesses, and extend methodologies is more valuable than pure theorem-proving ability for empirical work.
- Comfort with Uncertainty: Success requires being comfortable navigating the frontier of human knowledge where outcomes are unknowable.
- Alternative Routes:
- Research Engineering: Individuals with strong coding skills and ML understanding can contribute via industry roles (e.g., Google Brain residency) without a PhD.
- Policy and Strategy: Technical experts are needed in government and policy roles to inform regulations on autonomous weapons and AI strategy; "earning to give" is currently considered less impactful than filling the severe talent gap.
- Disagreements with Peers: Leicher notes a divergence with Dr. Dario Amodei (OpenAI) regarding research aptitude testing; Leicher emphasizes that true research potential is indicated by a natural obsession with unsolved puzzles and the ability to communicate complex ideas clearly, rather than just rapid model implementation.
- Market Reality: There is a critical shortage of qualified technical AI safety researchers; industry labs are eager to hire ML PhDs, who are highly employable even if they choose not to stay in AI safety.
- Upcoming Events: The Effective Altruism Global (EAGlobal) conference will take place in San Francisco on the second weekend of June; early bird tickets are available until March 18th.