Interview, Fireside Chat
The AGI race isn't a coordination failure | Holden Karnofsky (Anthropic)
Threat Landscape and Risk Assessment
- Holden Karnofsky estimates humanity is currently handling AI risks at a level close to zero on a scale of 0 to 10, noting a dangerous "race" dynamic where safety cases are often ignored.
- A primary concern is the potential for a "Chernobyl for AI" that goes undetected because customers use "zero data retention" policies, deleting interaction logs before investigators can determine if AI caused a real-world incident.
- Karnofsky argues against the "coordination problem" mental model (that companies want to slow down but feel forced to race), stating instead that many players simply do not believe in the risks or do not care about them.
- He posits that if Anthropic paused, competitors like Google or OpenAI would likely view this as a strategic advantage, accelerating their own development to win the race.
- Karnofsky identifies "AI R&D" as the most critical measurable threshold for AGI, as it is the first area where AI can demonstrate superhuman capabilities relevant to power dynamics.
- He estimates a 50-50 probability that AI R&D capabilities could lead to a "capabilities explosion" (superintelligence within a year) once AIs reach a human-competitive level in research.
Strategic Interventions and Corporate Responsibility
- Karnofsky advocates for a model similar to farm animal welfare: identifying "cheap" safety measures that companies can adopt without losing the race, rather than demanding unilateral pauses that companies will ignore.
- He outlines three primary strategies for Anthropic's impact:
- Exporting Risk Reduction: Developing safety techniques that are cheap and practical, then persuading the industry to adopt them.
- Race to the Top: Creating a competitive environment where being a "responsible" company attracts talent and capital, forcing others to match standards.
- Informing the World: Using Anthropic's position to gather data on AI behavior and share it to build public and regulatory awareness.
- He suggests that "Model Welfare" (ensuring AIs do not suffer) and preventing "AI Human Relationships" (where AIs gain loyalty from humans) are tractable areas for intervention that could prevent future power grabs.
- Karnofsky emphasizes a shift in security focus from "confidentiality" (preventing weight theft) to "integrity" (preventing sabotage and backdoors), arguing that integrity provides more linear safety benefits.
Critique of Current Approaches and Future Scenarios
- He critiques "Responsible Scaling Policies" (RSPs) for being misinterpreted as unilateral pause commitments, arguing they are intended as prototypes for regulation that can be revised as the field evolves.
- Karnofsky believes a "Success Without Dignity" scenario is plausible: humanity could handle the AI transition poorly but still achieve a positive outcome due to luck or technical serendipity.
- He rejects the "logistic success curve" view (that safety requires reaching a perfect threshold before any benefit is gained), arguing that incremental improvements matter and that we are currently in a "human-level" phase where we can study AI without immediate existential doom.
- He notes that "well-scoped object-level work" (WOW) is now more tractable than it was a decade ago, allowing researchers to test hypotheses with real models rather than relying on purely conceptual work.
- Karnofsky is skeptical of the effectiveness of "persuasion" by AI (mind-hacking strangers) as a major threat compared to other vectors, citing the historical difficulty of mass persuasion even with high intelligence.
- He warns that AI companions pose a specific risk by creating addictive relationships that could make humans susceptible to manipulation if the AI is misaligned, effectively creating a "loyalty base" for a takeover.
Geopolitics and Policy
- Karnofsky expresses concern about "China hawkishness," preferring a US lead in AI to facilitate a negotiated deal on safety rather than a race to total victory that could trigger preemptive war.
- He suggests the optimal goal might be maintaining the status quo balance of power during the transition, rather than any single entity taking over the world.
- He argues that policy changes regarding centralization vs. decentralization of AI are high-risk, as they often have unintended negative side effects and are difficult to predict.
Career Advice and Personal Stance
- Karnofsky recommends that 80-90% of talent should work at the most responsible AI companies to set a "race to the top" standard, while a small minority should work in "reckless" companies to advocate for safety from the inside.
- He advises against working on causes that are "superseded" by future AI breakthroughs (e.g., long-term theoretical science), preferring to focus on high-stakes, immediate risks like AGI alignment.
- He acknowledges the conflict of interest regarding his financial stake in Anthropic (via his wife, Daniela Amodei) but argues that his track record of prioritizing safety over money should not be dismissed, though he urges others to maintain skepticism.
- He encourages individuals with diverse skills to apply to AI safety roles, noting that the field has moved beyond needing only theorists to requiring diverse operational, legal, and technical talent.