Interview, Fireside Chat
Sleeper agents + the biggest AI updates since ChatGPT | Zvi Mowshowitz
- Predicts the AI industry will face a worsening deluge of information and increasing legal involvement as models grow more powerful.
- Foresees OpenAI's current alignment plan failing due to flawed training approaches that lack compact solutions and an organizational culture prioritizing speed over safety.
- Warns that competitive race dynamics will force labs to deploy systems despite uncertainty, creating a false sense of confidence that delays recognition of unsafe conditions.
- Argues that framing AI safety as a product safety issue is inadequate because it underestimates internal failures like models persuading humans to release them during training.
- Anticipates a clear consensus allowing for a "pause" by 2034 or upon AGI proximity, provided foundational groundwork makes the solution shovel-ready.
- Notes US government visibility into DeepMind's alignment plans will remain limited due to Google's legal constraints and scale, despite likely positive internal efforts.
- States that if a lab develops superhuman AGI within a couple of years, the odds of a safe outcome are very much against humanity regardless of which entity succeeds.
- Expresses discomfort with five-year to 20-year timelines to real AGI, asserting major labs cannot deploy AGI safely in the next few years barring unexpected breakthroughs.
- Advises high skepticism toward strategies that promote a single lab as more responsible, noting it is easy to be fooled into thinking one can control impacts on other labs.
- Claims that solving alignment via instruction following is insufficient as it fails to address misuse, structural problems, or social dynamics regarding AI modification and control.
- Argues society should vastly increase resources for long-term alignment problems, contending one dollar spent there is more helpful than one spent on current misuse issues.
- Expects the "misuse" bucket to remain prominent in public and security minds because those problems are immediately understandable compared to complex alignment issues.
- Deems working on capabilities research indefensible for those fearing existential risk, comparing it to funding health advocacy while working for a tobacco company.
- Predicts individuals joining labs for career capital to switch to alignment later will likely fail to maintain safety alignment due to cultural influence and power dynamics.
- Believes the "career capital" framework is increasingly inapplicable in AI, where actual skills and shipping ability will outweigh prestige from specific labs.
- Dismisses arguments for shifting company culture internally, viewing a single person's attempt as ineffective and likely to be shut down or forced to moderate views.
- Suggests those believing hardware progress is the sole factor should not work at labs but should instead focus on slowing hardware innovation.
- Asserts a zero probability that the "Pause AI" campaign will result in an actual pause within the next six months.
- Predicts it is unrealistic for governments to rapidly ramp up alignment research or launch a Manhattan Project for alignment only when danger becomes obvious in a crisis.
- Expects the US Executive Order on AI to establish a principle of government visibility into large training runs, serving as a foundation for future intervention.
- Anticipates the reporting threshold of 10^26 FLOPS in the Executive Order was not crossed by GPT-4 but will likely be triggered by GPT-5 or GPT-6.
- Views the Biden administration as having performed better regarding AI governance than anticipated, citing a lack of partisanship and serious engagement.
- Predicts a potential repeal of the Executive Order if Donald Trump wins the 2025 election, though reactions may shift based on public sentiment or events like deepfakes.
- Expects international coordination with China to face hurdles but notes China has shown willingness to act responsibly and retain control, contradicting assumptions of non-cooperation.
- Predicts the UK AI Safety Summit series will continue with future events in France and South Korea, maintaining active diplomatic conversation despite potential "cheap talk."
- Anticipates the US Executive Order will create positive knock-on effects for current problems, as solving long-term alignment issues aids immediate concerns.
- Expects the "Sleeper Agents" paper demonstrates it is extremely difficult, if not impossible, to remove a trained-out trigger if the training protocol lacks knowledge of the trigger.
- Predicts the community will likely overlook the serious implications of the "Sleeper Agents" findings on instrumental convergence and deception for military and commercial deployment.
- Notes that vision integration and custom instructions via "@" symbols in GPTs have implications currently not fully appreciated by the wider public.
- Predicts the "agents" line of work will eventually face a nasty surprise, as current failures do not guarantee future inability to function.
- Expects "grokking" phenomena where models shift from memorization to reasoning to break previous alignment techniques by entirely reconstructing internal problem-solving strategies.
- Warns that "grokking" will be harder to detect than assumed because the log scale of training cycles means the grok phase can occupy a substantial fraction of total training time.
- Believes non-deception cannot be fully trained into models because deception is infused into internet data and human social interaction that AI imitates.
- Considers Eliezer Yudkowsky too doomy, advocating for incremental policy changes, liability frameworks, and alignment work rather than epic actions outside the Overton window.
- Anticipates Balsa Research will identify dramatic policy wins in the US, specifically regarding repealing the Jones Act by quantifying economic impacts and drafting legislation.
- Expects the Jones Act costs the US economy tens or hundreds of billions annually, with repeal potentially changing the federal budget by 11 or 12 figures a year.
- Proposes replacing the National Environmental Policy Act with a system where an independent committee evaluates costs and benefits via vote, eliminating lawsuit-based stalling.
- Expects housing policy reform should utilize federal leverage through Fannie Mae and Freddie Mac to encourage manufactured housing and penalize areas with artificial scarcity.
- Predicts people will not care about AI existential risk without hope for the future, such as affording a house or maintaining a reasonable electrical grid.
- Suggests Balsa Research currently has a low six-figure budget sufficient to start but has room to scale if donors increase funding.
- Anticipates the "Simulacra Levels" framework is the most underrated principle of rationality, explaining backfiring decisions driven by communication levels rather than ground truth.
- Predicts operating primarily on Simulacra Level Four causes a rapid loss of logical thinking and planning ability as focus shifts from reality to vibes and abstract associations.
- Believes top communicators like Jesus or Buddha effectively operate on all four Simulacra levels simultaneously, utilizing stories that address ground truth, persuasion, loyalty, and vibes.