Interview
More accurately predicting the future | Philip Tetlock
Current Research Focus (Counterfactual Reasoning):
- Tetlock is currently investigating "backward reasoning" (counterfactuals) via the Civilization V forecasting tournament, where humans predict outcomes in simulated historical scenarios.
- The simulation allows for testing how accurately people can assess "what would have happened" if variables (e.g., leadership, economic conditions) were altered, addressing the "unknowable" nature of real-world history.
- Key Constraint: Machine learning is prohibited in the competition; participants must rely on human reasoning, mirroring the constraints faced by real-world intelligence analysts who cannot rerun history.
- Findings to Date: High proficiency in Civilization V does not guarantee high forecasting accuracy; strategic gameplay and causal forecasting skills are distinct competencies.
- Transferability Hypothesis: Tetlock posits that Civilization V mimics the real world regarding causal complexity, path dependency, and stochasticity, making it a viable proxy for training counterfactual reasoning skills applicable to geopolitics and economics.
Hybrid Forecasting and Algorithmic Performance:
- Algorithms struggle to outperform humans on low-data, high-ambiguity events (e.g., Syrian civil war, Brexit) where historical base rates are elusive.
- Machine learning succeeds in data-rich domains (e.g., macroeconomic statistics for OECD countries) but lacks traction in novel geopolitical scenarios.
- Simple Heuristics: Basic extrapolation algorithms (e.g., "predict no change" or "predict recent trend") often outperform human intuition, particularly for short-term forecasts where humans tend to over-extrapolate change.
- Superforecaster Traits: The strongest predictor of becoming a "superforecaster" is "perpetual beta" (a commitment to frequent, low-magnitude belief updating), which is roughly three times more predictive than raw intelligence (fluid or crystallized).
- Accuracy Boosts: Research indicates a ~40% accuracy gain from selecting superior forecasters, a ~10% gain from training, a ~10% gain from teaming, and a ~25% gain from algorithmic aggregation of the best forecasts.
Cognitive Biases and Probability Assessment:
- Binary Thinking: People without probabilistic training default to "0%," "50%," or "100%" probabilities, struggling to distinguish nuances in the "maybe" zone.
- Superforecaster Resolution: Top performers distinguish between 10–15 degrees of uncertainty, resisting the psychological urge to round to extremes.
- Political Tribalism: Expressing nuanced probabilities (e.g., 72% confidence in climate models) can alienate individuals from political tribes, creating social incentives to exaggerate certainty rather than accuracy.
- Risk Management: People often round low probabilities to 0% (ignoring tail risks) or overweight highly salient low-probability events (e.g., terrorism), leading to misallocated attention.
- Logical Consistency: Forecasts for extreme risks can be audited for logical consistency (e.g., ensuring a subset probability does not exceed the set probability) even when empirical verification is impossible.
Application to Personal Decision-Making:
- Career Forecasting (Academia): Base rates for academic success are low (often <5%), but broad averages can be misleading; success is highly concentrated at top-tier institutions where performance variance is significant.
- Self-Assessment: Individuals often lack accurate self-knowledge; the most reliable calibration often comes from a trusted mentor who provides frank feedback on specific capabilities.
- Business/Startups: Base rates for startup success are low even after VC screening; the optimal strategy involves pursuing high "fat-tail" opportunities (high upside, high failure rate) to capture outliers like Facebook or Google.
- Deliberation vs. Implementation: Forecasting accuracy is vital during the "deliberation" phase of decision-making, but once committed, individuals often shift to an "implementation mindset" requiring optimism to mobilize others and sustain effort.
Institutional and Systemic Barriers:
- Status Threat: Forecasting tournaments threaten existing status hierarchies (e.g., senior analysts vs. young forecasters), leading to resistance from established institutions like the intelligence community and media outlets.
- Verbal Land vs. Quantification: Experts and politicians often use vague language ("distinct possibility") to mask uncertainty; quantifying these claims reveals wide calibration gaps (20–80% range).
- Training Efficacy: The "Champs Nul" training module provided a 6–12% performance boost in tournaments, though the transferability of specific calibration tools (e.g., for math or polls) to complex real-world domains remains an open research question.
- Future Trajectory: Tetlock predicts a slow, halting adoption of forecasting methods in government and business by 2030–2040, driven by a growing awareness of cognitive limitations and the need to mitigate "bait-and-switch" heuristics (substituting a difficult question with an easier status-based one).
Long-Term Forecasting and Uncertainty:
- Limits of Prediction: Superforecasters lose their advantage over chance roughly 5–10 years out, and completely disappear as a predictive advantage at the century scale due to compounding randomness (the "card shuffling" of history).
- Butterfly Effects: While the world is chaotic and individuals (e.g., Hitler, Stalin) can alter historical trajectories, predicting such "black swan" events remains impossible; historical contingency is high.
- Extremizing Strategy: Aggregating forecasts with "extremizing" algorithms (pushing predictions toward the poles) worked well in 2011–2015 but is risky in volatile environments; its utility depends on the specific accuracy function and cost of errors.
- Prediction Markets vs. Tournaments: Forecasting tournaments currently outperform prediction markets when the latter lack depth and liquidity, but deep, liquid markets may theoretically outperform if given sufficient time and data.
Tools and Resources Mentioned:
- Civilization V Sign-up: 80k.link/civ for the counterfactual forecasting tournament.
- Calibration Training: 80000hours.org/calibration-training tool funded by Open Philanthropy.
- Reference Materials: A 20-minute summary by Daniel Cocotilo for AI Impacts; Spiro Makridakis's M4 forecasting competition.
- Bridgewater Associates: Cited as a model for "accuracy games" and radical transparency, though with high human costs.