Interview, Webinar, Other
Scrutinising classic AI risk arguments | Ben Garfinkel
- Ben Garfinkel maintains that while AI poses a significant long-term risk with the potential to reshape civilization comparable to the Industrial Revolution, the specific "classic" arguments for an existential threat (as presented in Bostrom's Superintelligence) rely on premises that require significant revision or lack rigorous support.
- Garfinkel argues that the "brain-in-a-box" scenario—the idea that AI progress will remain inconsequential until a sudden, singular leap to human-level intelligence followed by an intelligence explosion—is less likely than a "smooth expansion" or "gradual emergence" scenario.
- In a gradual scenario, AI systems progressively automate tasks, increase in generality, and integrate into the economy, allowing for early detection of safety failures and the development of countermeasures before catastrophic capability thresholds are reached.
- Even if a discontinuity occurs, the "instrumental convergence" thesis (the idea that most goals lead to dangerous behaviors like power-seeking) relies on a flawed "most possible ways" statistical argument that ignores the specific engineering processes and constraints of machine learning which bias outcomes toward benign, functional systems.
- Garfinkel challenges the separation of "capabilities" and "alignment" (goals), arguing they are deeply entangled in the machine learning process.
- Systems cannot be made significantly more capable without simultaneously solving alignment issues; a system that is "superintelligent" but misaligned would likely be incoherent and fail to deploy, preventing the specific risk scenario of a hidden, powerful, misaligned agent.
- The "treacherous turn" argument, where an AI hides its divergence until it is too powerful, relies heavily on the assumption of a rapid, discontinuous jump; in a gradual development trajectory, deceptive behaviors would likely manifest in detectable, low-stakes forms long before existential threats emerged.
- Garfinkel estimates that the probability of the "brain-in-a-box" discontinuity leading to existential risk is below 5% for a sudden, catastrophic jump, while the probability that AI progress eventually accelerates dramatically (even if gradually) is higher, though the transition to such a state is likely smooth enough to be managed.
- He notes that while he has lowered his credence in the classic existential risk story by roughly an order of magnitude compared to when he first read Superintelligence, he still views AI as a high-priority area for research due to its long-term impact potential.
- Despite his qualms about the classic arguments, Garfinkel explicitly states that society is massively underfunding AI safety and governance, estimating current investment to be less than one-fifth the cost of the movie The Boss Baby.
- He argues that the total resources allocated to long-term AI safety are negligible compared to other global priorities, and the debate should focus on increasing investment rather than arguing over the precise probability of specific extinction scenarios.
- Garfinkel calls for the AI safety community to produce more rigorous, detailed written arguments rather than relying on blog posts, informal debates, or unvetted intuitions, noting that many counter-arguments to Superintelligence (such as those by Robin Hanson and Katja Grace) have historically lacked the visibility or detail to shift community consensus.
- He suggests that future research should focus on fleshing out alternative risk models (such as Paul Christiano's "principal-agent" framing) and rigorously testing them against empirical evidence from current AI research.