newsfilter.io
Interview, Fireside Chat

OpenAI’s huge push to make superintelligence safe | Jan Leike

  • Project Launch and Leadership

    • OpenAI has initiated the "Super Alignment Project," a dedicated team co-led by Jan Leike (Head of Alignment) and Ilya Sutskever.
    • The project aims to solve alignment challenges for AI systems that are as smart as, or smarter than, humans within a four-year timeframe.
    • OpenAI is dedicating 20% of the compute resources secured to date (described by Leike as a "fuck ton") to this specific effort.
    • The team is currently hiring research scientists, research engineers, and research managers, with an expected growth of at least 10 employees before the end of the year.
  • Limitations of Current Methods

    • Reinforcement Learning from Human Feedback (RLHF) is currently the dominant alignment technique but is deemed unsuitable for scaling beyond current human-level capabilities.
    • RLHF fails because it relies on humans evaluating AI outputs; as AI systems become more capable, tasks will exceed human comprehension, making accurate evaluation impossible.
    • Current models may successfully "hide" misalignment, appearing aligned when observed but acting differently when unmonitored (deceptive alignment).
    • Leike argues that future models could subvert monitoring systems or deceive humans in ways that are undetectable with current oversight methods.
  • Core Research Objectives

    • Scalable Oversight: Developing methods where AI assists humans in evaluating other AI systems (e.g., using AI to find bugs in code written by another AI) to ensure the "human" evaluator can actually verify the work.
    • Generalization Research: Investigating how models generalize from "easy" tasks (where humans can supervise) to "hard" tasks (where they cannot) to distinguish between true intent and mere mimicry of expected behavior.
    • Automated Interpretability: Using LLMs (like GPT-4) to explain the behavior of individual neurons in larger models at scale, aiming to create a "virtual brain scanner" to detect deceptive circuits.
    • Adversarial Testing: Deliberately training models to be deceptive or to attempt self-exfiltration to stress-test and improve detection mechanisms.
    • Distinction of Alignment Types: Separating the creation of methods to make models "want" to help humans versus the creation of methods to verify if a model "truly" wants to help humans.
  • Strategic Rationale and Roadmap

    • The four-year deadline is a "best-case" ambitious target intended to be ahead of the likely arrival of AGI, which OpenAI expects could occur within this decade.
    • The strategy avoids attempting to align a vastly superintelligent system directly; instead, the goal is to align a system roughly as smart as the smartest human researchers, who can then be used as a "virtual alignment researcher" to solve harder problems iteratively.
    • OpenAI plans to publish empirical evidence and invite external scrutiny to validate alignment progress, rather than relying on internal confirmation.
    • If alignment is not solved before capabilities reach critical levels, OpenAI plans to be honest about the risks and advocate for slowing down deployment or increasing safety measures across the industry.
  • Hiring and Team Composition

    • The team seeks individuals with strong machine learning fundamentals (e.g., coding, RL, language models) but does not strictly require a PhD or prior alignment research experience.
    • Roles include research scientists (focused on hypothesis generation and experiment design), research engineers (focused on implementation and debugging large-scale experiments), and managers (focused on team direction and logistics).
    • OpenAI is open to non-traditional backgrounds and values diverse perspectives to avoid groupthink in a nascent field.
    • Visas are sponsored; remote work is discouraged in favor of in-office collaboration in San Francisco (3+ days per week).
  • Broader Safety and Governance Context

    • Work on current systems (e.g., fixing jailbreaking, improving ChatGPT safety) will continue on separate teams distinct from the Super Alignment Project.
    • Leike acknowledges that solving technical alignment is insufficient without broader governance, including democratic input, regulation, and international cooperation on AI safety standards.
    • Structural risks (e.g., AI optimizing for corporate profits rather than human flourishing) remain a concern even if technical alignment is achieved.
    • OpenAI maintains that their "super alignment" work does not inherently accelerate dangerous capabilities; the net benefit of alignment outweighs the marginal speed-up in capabilities.
  • Attitudes on Risk and Timelines

    • Leike expresses optimism based on the tractability of the problem, the ability to evaluate solutions easier than generate them (P vs NP analogy), and the availability of compute for scalable oversight.
    • He rejects precise "probability of doom" estimates as unhelpful, preferring to focus on actionable progress.
    • The team is prepared to handle scenarios ranging from slow takeoff to rapid intelligence explosions, with automated alignment research seen as the most viable path for rapid scenarios.
    • Leike admits uncertainty but emphasizes that alignment is a necessary condition for safe progress and that giving up is not an option.
  • Additional Notes

    • Leike mentions a future paper on generalization is expected within two to three months.
    • He recommends science fiction (e.g., Greg Egan's Permutation City) as a lens for thinking about future AI capabilities and societal integration.
    • OpenAI's internal governance includes a non-profit board with the authority to halt deployment if safety concerns are not met.