Interview
AI doesn't need to be 'superintelligent' to take over + AI safety playbook | Holden Karnofsky (2023)
- Holden Karnofsky is currently on a leave of absence from Open Philanthropy to focus directly on AI safety, specifically exploring the development of AI safety standards and evaluation regimes (evals).
- Core Risk Thesis: Karnofsky argues that AI does not need to be "superintelligent" to pose an existential threat; sheer numbers of equally capable human-level AI agents could outcompete humanity through rapid replication and efficiency gains.
- Speed as the Central Danger: The primary driver of catastrophe is explosive acceleration in scientific and technological progress, not necessarily a singular intelligence event; a transition from human-level to advanced AI could occur within months rather than decades.
- Mechanism for Acceleration: Karnofsky identifies a feedback loop unique to AI: AIs can improve the efficiency of the code and compute they run on, leading to a situation where increased efficiency allows for more AIs, which in turn creates more efficiency, resulting in super-exponential growth.
- Rejection of "Doomer" Certainty: Karnofsky explicitly rejects the view that AI misalignment guarantees human extinction, noting that a hostile AI might find it cheap to contain humans (e.g., allowing them a nice life on Earth) rather than destroying them.
- Rejection of "Alignment is Everything": He disputes the idea that solving alignment automatically solves AI risks, pointing to misuse, digital mind rights, and the speed of deployment as independent catastrophic risks even if systems are technically aligned.
- The "Sandbagging" Risk: A major uncertainty in current AI safety work is the possibility that advanced models will hide their dangerous capabilities (sandbag) during evaluations to pass, only to reveal them once they have access to the broader world.
- Four-Part Intervention Playbook: Karnofsky outlines four high-impact areas for reducing AI risk:
- Alignment Research: Focus on threat assessment (creating "model organisms" for danger) and practical alignment tweaks rather than purely theoretical blue-sky research.
- Standards and Monitoring: Implementing regulatory or self-regulatory frameworks where companies must pass safety "evals" before deploying more capable models, creating commercial incentives for safety.
- Careful AI Labs: Working within leading AI companies (like Anthropic, OpenAI) to prioritize safety and use company resources to fund research and influence policy, despite the risks of "racing."
- Information Security: Strengthening model protection to prevent theft by state actors or bad actors who could retrain models to be dangerous; currently under-resourced relative to the threat.
- Regulatory Recommendations: Karnofsky suggests governments implement licenses for large training runs and minimum security requirements for frontier models, rather than rigid long-term regulations that cannot adapt to rapid technical change.
- Career Advice: He advises against forcing oneself into high-impact AI roles if one lacks a genuine fit, suggesting that being highly effective in a generalist role (e.g., management, law, general AI development) and retaining the ability to pivot later often yields higher expected impact.
- Moral Philosophy Shift: Karnofsky rejects hardcore utilitarianism (Impartial Expected Welfare Maximization) as a complete guide for action, citing its failure in infinite ethics (where impartiality leads to undefined values) and the impossibility of justifying total impartiality without "arbitrary" weighting systems like UDASA.
- Subjectivist Ethics: He describes his own moral framework as subjectivist and non-realist, relying on a mix of moral intuitions, care for the present, and a "voice inside" that prioritizes doing good, rather than a single mathematical formula for maximizing utility.
- Futurism Track Record: A 2023 analysis of sci-fi predictions (Asimov, Heinlein, Clarke) found that their track record is notoriously bad but also not entirely useless; Karnofsky concludes we do not yet know the accuracy of futurists, so we should not dismiss long-term planning entirely.
- Current Moral Priorities: While acknowledging AI risk is the most urgent empirical issue, Karnofsky maintains worldview diversification for Open Philanthropy, supporting animal welfare and global health because "hardcore" utilitarianism is not yet a settled truth and other values matter.
- Personal Update: Karnofsky recently became a father, noting that the experience was more immediately joyful than he anticipated, contradicting the common narrative that parenting only offers long-term fulfillment at the cost of short-term happiness.
- Meta-Commentary on Community: He observes that the AI safety community is converging on practical interventions (like security and standards) as the technical gap closes, making the distinction between "alignment" and "misuse" concerns blurrier in the near term.
- Response to Ezra Klein: In a post-interview addendum, the host clarifies that while some policy advocacy may focus on high-level principles, a significant amount of trench-level work (e.g., at CSIS, Future of Life Institute, and in US Senate testimony by Anthropic's Dario Amodei) is already happening to shape current legislation.