newsfilter.io
Interview, Fireside Chat

#23 - How to actually become an AI alignment researcher, according to Dr Jan Leike

  • Rapid improvement in AI and machine learning is expected to continue, enabling progress on global issues such as poverty and animal suffering.
  • Current plans prioritize technical approaches to AGI safety, specifically advancing deep reinforcement learning from human preferences to define objective functions through simple feedback mechanisms like upvoting.
  • Future development aims to create systems where non-experts can teach agents arbitrary objectives, requiring better feedback methods beyond clip comparisons and more stable deep reinforcement learning algorithms that reduce performance variance across random seeds.
  • Machine learning security is identified as a critical challenge where minimal input perturbations can drastically alter outputs, with black box attacks performing surprisingly well and current defenses remaining ineffective.
  • The field of AI safety is characterized as nascent with significant "low-hanging fruit," yet faces a severe talent gap that hinders DeepMind's ability to hire researchers despite active recruitment efforts.
  • Specific technical problem areas noted for further exploration include anomaly detection, safe exploration, maximizing reward with implicit side-effect regularization, and determining the optimal volume of human feedback required over time.
  • There is a strong recommendation for technical experts to enter policy, government, or strategy roles, alongside a belief that a machine learning PhD is the primary path for skill building in this domain.
  • Long-term predictions indicate that most technical AI safety researchers will eventually utilize machine learning approaches, though the speaker plans to prioritize theoretical methods despite current difficulty in execution.
  • Collaboration opportunities are anticipated with OpenAI and DeepMind, with a specific hope for deeper alignment between these entities to address shared safety challenges.
  • Career opportunities are highlighted for PhDs in machine learning seeking applied impact, with specific mentions of health-related projects at DeepMind and the broader sector.