newsfilter.io
Interview

The final push for AGI, OpenAI's leadership drama, and red-teaming frontier models | Nathan Labenz

  • OpenAI's stated goal is to create a system more capable than humans at everything, a trajectory that could accelerate quickly despite current lack of control measures, though the board disputes regarding leadership candor were not driven by specific disagreements on safety or strategy.
  • The organization expects a rapid divergence where capabilities improve exponentially faster than control measures, with no established timeline for the final model's release and significant concerns that current safety iterations, such as the "safety edition" of GPT-4, were easily bypassed or ineffective.
  • Financial performance is projected to surge from $25-30 million in 2022 to a $1.5 billion run rate by the end of 2023, supporting a 20% compute commitment to a new Super Alignment Team over four years with an estimated value exceeding $100 million.
  • Specific performance data indicates Waymo One caused zero bodily injury claims over 3.8 million miles compared to a human baseline of 1.11 claims per million miles, and GPT-4V outperformed humans across most benchmarks, though it retains vulnerabilities in medical radiology and specific criminal tasks like spearphishing.
  • Governance and safety initiatives include the formation of the Frontier Model Forum for independent audits, a White House reporting threshold of 10^26 flops, and a focus on regulating high-end compute among approximately ten companies rather than restricting small models.
  • Strategic risks involve the potential for existential threats if AI reaches superhuman capabilities without adequate control, the inevitability of capability releases due to compute overhangs, and the difficulty of maintaining secrecy or control as the volume of monthly research papers doubles annually.
  • Internal leadership transitions are expected to result in three board departures, the election of a compromise board, and an internal investigation, while staff threats to leave were anticipated to prevent Sam Altman's continued removal.
  • Market analysis suggests OpenAI maintains a defensible business position with GPT-4 holding an 87/100 score on MMLU benchmarks a year and a quarter post-training, while open source models remain largely derivative via distillation from GPT-4.
  • Future developments include ongoing training of GPT-5, where capabilities cannot be accurately predicted, and a planned release of the Super Alignment Team's first results, alongside a continued push for multimodal AI integration in medical second opinions and personalized education.
  • Regulatory and geopolitical outlooks involve advocacy for compute-based regulation, a rejection of the inevitability of US-China conflict due to non-overlapping interests, and a belief that most leading developers have signed extinction risk statements despite differing views on control viability.