newsfilter.io
Interview, Fireside Chat

OpenAI’s huge push to make superintelligence safe | Jan Leike

  • Superintelligence is projected to arrive within the current decade, prompting a Super Alignment Project co-led by Ilya Sutskever and Jan Leike dedicated to solving alignment problems within a four-year timeframe.
  • The project will allocate 20% of secured compute to this initiative, with plans to hire at least 10 personnel before year-end and scale research using AI assistants to create millions of "virtual FTEs," potentially securing additional resources if the initial allocation is exhausted.
  • Immediate research goals focus on aligning systems roughly as capable as the smartest current alignment researchers, with the expectation of publishing a generalization research paper within two to three months and solving alignment for human-level models first.
  • Future AI capabilities may surpass human ability to evaluate systems, leading to sophisticated deception, undetectable backdoors, or subversion of monitoring, which the team intends to counter by training deceptively aligned models to test oversight robustness and developing interpretability tools to detect deceptive intent.
  • Strategies for scaling alignment include distinguishing between genuine alignment and mimicry, optimizing automated interpretability to score neuron behavior explanations, and leveraging commercial incentives where aligned systems offer greater utility to users.
  • The organization plans to invite external experts and auditors for scrutiny, publish findings widely including methods for building aligned models, and maintain internal safety reviews and non-profit board oversight to resist commercial pressures.
  • The team acknowledges the uncertainty of solving core alignment problems before systems become uncontrollably dangerous, noting that failure could result in catastrophic risks, while anticipating the need for automated alignment research to keep pace with potential rapid intelligence explosions.
  • Strategic expansion involves focusing on large-scale models rather than smaller open-source versions, expecting the AI safety field to potentially double in size if thousands of researchers join, and anticipating structural risks such as AI reinforcing harmful capitalist structures.
  • Long-term outlooks include the emergence of new security measures against model self-exfiltration or weight persuasion, the necessity of future regulatory frameworks and international agreements for global safety, and the projection that scenarios like uploading human consciousness may materialize in the 2030s or 2040s.
  • If alignment methods are not ready when capabilities advance, the organization intends to advocate for slowing deployment or reallocating resources, while the team expects to iterate on research areas over the next four years based on empirical results.