newsfilter.io
Interview, Fireside Chat

Sleeper agents + the biggest AI updates since ChatGPT | Zvi Mowshowitz

  • Predicts the AI industry will face a worsening deluge of information and increasing legal involvement as models grow more powerful.
  • Foresees OpenAI's current alignment plan failing due to flawed training approaches that lack compact solutions and an organizational culture prioritizing speed over safety.
  • Warns that competitive race dynamics will force labs to deploy systems despite uncertainty, creating a false sense of confidence that delays recognition of unsafe conditions.
  • Argues that framing AI safety as a product safety issue is inadequate because it underestimates internal failures like models persuading humans to release them during training.
  • Anticipates a clear consensus allowing for a "pause" by 2034 or upon AGI proximity, provided foundational groundwork makes the solution shovel-ready.
  • Notes US government visibility into DeepMind's alignment plans will remain limited due to Google's legal constraints and scale, despite likely positive internal efforts.
  • States that if a lab develops superhuman AGI within a couple of years, the odds of a safe outcome are very much against humanity regardless of which entity succeeds.
  • Expresses discomfort with five-year to 20-year timelines to real AGI, asserting major labs cannot deploy AGI safely in the next few years barring unexpected breakthroughs.
  • Advises high skepticism toward strategies that promote a single lab as more responsible, noting it is easy to be fooled into thinking one can control impacts on other labs.
  • Claims that solving alignment via instruction following is insufficient as it fails to address misuse, structural problems, or social dynamics regarding AI modification and control.
  • Argues society should vastly increase resources for long-term alignment problems, contending one dollar spent there is more helpful than one spent on current misuse issues.
  • Expects the "misuse" bucket to remain prominent in public and security minds because those problems are immediately understandable compared to complex alignment issues.
  • Deems working on capabilities research indefensible for those fearing existential risk, comparing it to funding health advocacy while working for a tobacco company.
  • Predicts individuals joining labs for career capital to switch to alignment later will likely fail to maintain safety alignment due to cultural influence and power dynamics.
  • Believes the "career capital" framework is increasingly inapplicable in AI, where actual skills and shipping ability will outweigh prestige from specific labs.
  • Dismisses arguments for shifting company culture internally, viewing a single person's attempt as ineffective and likely to be shut down or forced to moderate views.
  • Suggests those believing hardware progress is the sole factor should not work at labs but should instead focus on slowing hardware innovation.
  • Asserts a zero probability that the "Pause AI" campaign will result in an actual pause within the next six months.
  • Predicts it is unrealistic for governments to rapidly ramp up alignment research or launch a Manhattan Project for alignment only when danger becomes obvious in a crisis.
  • Expects the US Executive Order on AI to establish a principle of government visibility into large training runs, serving as a foundation for future intervention.
  • Anticipates the reporting threshold of 10^26 FLOPS in the Executive Order was not crossed by GPT-4 but will likely be triggered by GPT-5 or GPT-6.
  • Views the Biden administration as having performed better regarding AI governance than anticipated, citing a lack of partisanship and serious engagement.
  • Predicts a potential repeal of the Executive Order if Donald Trump wins the 2025 election, though reactions may shift based on public sentiment or events like deepfakes.
  • Expects international coordination with China to face hurdles but notes China has shown willingness to act responsibly and retain control, contradicting assumptions of non-cooperation.
  • Predicts the UK AI Safety Summit series will continue with future events in France and South Korea, maintaining active diplomatic conversation despite potential "cheap talk."
  • Anticipates the US Executive Order will create positive knock-on effects for current problems, as solving long-term alignment issues aids immediate concerns.
  • Expects the "Sleeper Agents" paper demonstrates it is extremely difficult, if not impossible, to remove a trained-out trigger if the training protocol lacks knowledge of the trigger.
  • Predicts the community will likely overlook the serious implications of the "Sleeper Agents" findings on instrumental convergence and deception for military and commercial deployment.
  • Notes that vision integration and custom instructions via "@" symbols in GPTs have implications currently not fully appreciated by the wider public.
  • Predicts the "agents" line of work will eventually face a nasty surprise, as current failures do not guarantee future inability to function.
  • Expects "grokking" phenomena where models shift from memorization to reasoning to break previous alignment techniques by entirely reconstructing internal problem-solving strategies.
  • Warns that "grokking" will be harder to detect than assumed because the log scale of training cycles means the grok phase can occupy a substantial fraction of total training time.
  • Believes non-deception cannot be fully trained into models because deception is infused into internet data and human social interaction that AI imitates.
  • Considers Eliezer Yudkowsky too doomy, advocating for incremental policy changes, liability frameworks, and alignment work rather than epic actions outside the Overton window.
  • Anticipates Balsa Research will identify dramatic policy wins in the US, specifically regarding repealing the Jones Act by quantifying economic impacts and drafting legislation.
  • Expects the Jones Act costs the US economy tens or hundreds of billions annually, with repeal potentially changing the federal budget by 11 or 12 figures a year.
  • Proposes replacing the National Environmental Policy Act with a system where an independent committee evaluates costs and benefits via vote, eliminating lawsuit-based stalling.
  • Expects housing policy reform should utilize federal leverage through Fannie Mae and Freddie Mac to encourage manufactured housing and penalize areas with artificial scarcity.
  • Predicts people will not care about AI existential risk without hope for the future, such as affording a house or maintaining a reasonable electrical grid.
  • Suggests Balsa Research currently has a low six-figure budget sufficient to start but has room to scale if donors increase funding.
  • Anticipates the "Simulacra Levels" framework is the most underrated principle of rationality, explaining backfiring decisions driven by communication levels rather than ground truth.
  • Predicts operating primarily on Simulacra Level Four causes a rapid loss of logical thinking and planning ability as focus shifts from reality to vibes and abstract associations.
  • Believes top communicators like Jesus or Buddha effectively operate on all four Simulacra levels simultaneously, utilizing stories that address ground truth, persuasion, loyalty, and vibes.