Interview, Podcast
What superforecasters and experts think about existential risks | Ezra Karger
- Definition of "Bad Outcomes Cluster": The study defined a cluster of catastrophic outcomes to include AI-caused human extinction, or cases where AI misuse/misalignment causes a population drop of 50%+ and a significant decline in human well-being.
- AI risk skeptics assigned a 30% probability that one of these outcomes would occur within the next 1,000 years.
- AI risk concern groups assigned a 40% probability to the same cluster of outcomes over the same 1,000-year horizon.
- Disagreement primarily centers on the timing of these risks rather than their long-term existence; skeptics view risks as occurring further in the future compared to concern groups.
- Existential Risk Persuasion Tournament (XPT) Overview: The Forecasting Research Institute (FRI) conducted a tournament with over 100 domain experts, 100 superforecasters, and public participants to forecast catastrophic risks (nuclear, AI, bio, climate) by 2100 and 1,000 years.
- Methodology: The process involved four stages: individual initial forecasts, team-based collaboration within homogeneous groups, mixed-team collaboration, and cross-team review, generating over 5 million words of debate rationales.
- Key Finding on Definitions: Prior to the tournament, estimates varied widely due to undefined terms; the project required participants to define "extinction" and "catastrophe" (e.g., 10% population loss in 5 years) to ensure comparable data.
- Headline Probability Estimates (by 2100):
- Total Extinction Risk: Domain experts estimated a 6% chance of human extinction; superforecasters estimated a 1% chance.
- Catastrophic Risk (10%+ population loss): Experts estimated a 20% chance; superforecasters estimated a 9% chance.
- AI-Specific Extinction Risk: Domain experts gave a 3% probability; superforecasters gave a 0.38% probability.
- Nuclear Extinction Risk: Domain experts gave a 0.5% probability; superforecasters gave a 0.07% probability.
- Non-Anthropogenic Risks (e.g., asteroids): Both groups agreed closely, with experts estimating 0.004% and superforecasters giving virtually identical forecasts, suggesting experts are not universally over-estimating risk.
- AI Capability and Economic Forecasts:
- Timeline Agreement: Both groups anticipated advanced AI systems would emerge soon, with experts predicting 2046 and superforecasters predicting 2060.
- Short-term Capability Milestones: Experts predicted AI would write three NYT bestsellers by 2038; superforecasters predicted 2050.
- Economic Growth Disagreement: Experts assigned a 25% chance that annual global GDP growth would increase by >15% year-over-year before 2100 due to AI; superforecasters assigned only a 2-3% chance.
- Compute Spend Forecasts (2024): Superforecasters predicted $35 million spent on the largest AI run; experts predicted $65 million.
- AI Adversarial Collaboration (Follow-up Study): A smaller study brought 11 AI skeptics and 11 AI concerners together to test four hypotheses for why they disagreed.
- Rejection of Hypotheses 1 & 2: Disagreement was not due to low engagement quality, nor was it fully explained by short-term expectations (short-term cruxes could only explain ~1% of the 25 percentage point belief gap).
- Support for Hypothesis 3 (Long-term Expectations): By 2100, both groups agreed powerful AI would exist (88-90% probability), but differed on outcomes; skeptics anticipated lower median human well-being even if extinction is avoided, while concern groups anticipated higher well-being if no extinction occurs.
- Support for Hypothesis 4 (Worldview Disagreements): Skeptics viewed the world as continuous and resilient (slow-moving change), while concern groups viewed it as discontinuous (fast-moving, potentially catastrophic shifts).
- Convergence Results: Little convergence occurred during the 8-week debate; skeptic beliefs shifted from 0.1% to 0.12%, while concern beliefs shifted from 25% to 20%, with changes largely attributed to external world events (e.g., May 2023 AI developments) rather than the debate itself.
- Top Value-of-Information Cruxes:
- For concern groups: A major powers war by 2030 and an independent body (e.g., Meter/Arch Eval) confirming AI can autonomously replicate/resources/evade deactivation.
- For skeptic groups: Superforecasters changing their minds on extinction risk and the development of lethal technologies capable of human extinction.
- Convergence Crux: An independent body confirming AI's ability to evade deactivation and acquire resources was the single factor most likely to cause both groups to converge.
- Public Elicitation and Calibration:
- Initial public estimates for AI extinction risk (2%) fell between expert and superforecaster figures.
- When given access to reference classes (low-probability events like lightning strikes), public estimates for extinction risk dropped drastically to 1 in 15 million (total) and 1 in 30 million (AI).
- This suggests public forecasts are highly sensitive to framing and reference classes, highlighting a "low probability miscalibration" problem.
- Forecasting Science and Methodology:
- Incentives: The project used proper scoring rules for short-term questions and "intersubjective metrics" (predicting what other groups would say) for long-term questions where ground truth is unknowable.
- LLM Augmentation: A working paper indicates that providing humans with Large Language Models improves forecast accuracy, though risks of misinformation exist.
- LLM Performance: Current AI-based forecasting systems generally lag behind human accuracy but are approaching human performance on questions with high uncertainty (50% probability).
- Future Research FRI is hiring to conduct large-scale experiments on eliciting low-probability forecasts and testing intersubjective metrics.
- Critiques and Limitations:
- Attrition: High dropout rates among experts complicated longitudinal analysis, though final forecasts were deemed representative.
- Sample Representativeness: Experts were a convenient sample (500 applicants, not random) with a bias toward the effective altruism community; future studies aim for more defined recruitment baselines.
- Engagement Quality: While most engagement was high, some conversations were deemed low quality; the project concluded that structured adversarial collaboration (e.g., facilitated debates) may be necessary for deeper convergence.
- Correlation of Risks: Beliefs about nuclear, bio, and AI risks were highly correlated, suggesting an underlying worldview factor (e.g., perceived fragility of the world) rather than independent domain-specific assessments.
- Recommendations:
- FRI released an 800-page report intended as a reference for policymakers and academics; an executive summary is available for quick reference.
- Ezra Karger recommends Moving Mars (Greg Bear) for geopolitical tech scenarios, The Second Kind of Impossible (Paul Steinhardt) for scientific discovery, and The Rise and Fall of American Growth (Robert Gordon) for historical economic context.