Interview
#15 - Prof Tetlock on how chimps beat Berkeley undergrads and when it’s wise to defer to the wise
Research Foundation & Scoring:
- Philip Tetlock's 35-year research program includes tens of thousands of participants and millions of predictions, primarily scored using the Brier score system developed by meteorologist Glenn Breyer.
- Lower Brier scores indicate higher accuracy; a perfect score of 0 implies clairvoyance, while a score of 2 indicates "inverse clairvoyance" (perfectly wrong predictions).
- A random-forecasting baseline (equivalent to a dart-throwing chimp or coin flip) yields a score of 0.5.
- Initial findings revealed that political experts performed only marginally better than dart-throwing chimps, with experts failing to significantly outperform "attentive readers" of elite press (e.g., New York Times) and losing to simple extrapolation algorithms.
- Experts significantly outperformed only one group: Berkeley undergraduates, who performed worse than chance due to extreme overconfidence and lack of self-awareness.
Key Behavioral Findings (Hedgehogs vs. Foxes):
- "Hedgehogs" (those with one big theory) performed worse than "foxes" (those who stick to data and acknowledge multiple viewpoints), particularly in long-range forecasts.
- Hedgehogs exhibited higher overconfidence, claiming events were 100% certain when they occurred 80% of the time, and events predicted at 80% occurred only 65% of the time.
- Better forecasters displayed curiosity, open-mindedness, and a high tolerance for cognitive dissonance, allowing them to entertain counterfactuals that challenged their ideological priors.
- Poor performers often relied on post-hoc defenses like "timing errors" or "exogenous shocks" rather than updating beliefs after failed predictions.
The "Superforecasters" & Aggregation:
- The later Superforecasting research identified a subset of individuals who consistently made highly accurate predictions through rigorous training and teamwork.
- Aggregating judgments from superforecasters produced results superior to those of intelligence community analysts with access to classified data.
- Superforecasting teams utilized a weighted average of the best recent forecasts, "extremizing" the result based on the diversity of viewpoints among forecasters (cognitive diversity).
- Extreme caution is advised when extremizing judgments within groups that have already collaborated extensively, as they risk "self-extremizing" rather than reflecting true cognitive diversity.
New Tournament (Hybrid Forecasting Competition):
- A new IARPA-sponsored tournament pits humans, machines, and human-machine hybrids against each other to determine optimal forecasting methods.
- The competition targets two domains: quantitative problems (where AI currently holds an advantage) and idiosyncratic, context-specific questions (e.g., Yasser Arafat's autopsy, Syrian civil war duration) where human intuition may still excel.
- Volunteers can participate at hybridforecasting.com, where they are assigned to specific forecasting algorithms to test different aggregation strategies.
Calibration, Training, and Bias:
- Calibration training shows limited transferability; skills learned in one domain (e.g., poker) do not automatically generalize to unrelated domains (e.g., geopolitics) unless the training includes understanding the underlying metrics of probability.
- Tetlock's personal experience in a tournament demonstrated that "tweaking" optimal algorithmic predictions led to a drop from 2nd place to 35th, underscoring the value of deferring to established aggregation logic over individual cognitive effort.
- Publicly committing to a strong position can freeze beliefs and trigger "defensive bolstering," making it harder to update views in the face of new evidence.
Forecasting Extreme Events ("The Tails"):
- Assessing the accuracy of forecasts for low-probability, high-impact events (e.g., nuclear war, pandemics) is difficult due to the scarcity of actual events for validation.
- Validity in the "tails" is often assessed through logical coherence checks, such as verifying "temporal scope sensitivity" (ensuring probabilities increase appropriately as the time horizon lengthens).
- "Question clustering" is proposed as a method to predict long-term risks by aggregating answers to specific, diagnostic intermediate questions (e.g., driverless car adoption rates as a proxy for the Fourth Industrial Revolution).
Institutional & Societal Implications:
- Tetlock opposes the "populist know-nothingism" that dismisses all expertise, arguing that ignoring experts is dangerous, though justified skepticism of overconfident pundits is necessary.
- He distinguishes between domains where expertise should be highly deferred (e.g., surgery, where feedback is rapid and clear) versus macro-political domains where feedback is delayed and ambiguous.
- The "Alpha Pundit Challenge" is a proposed initiative to extract implicit forecasts from media commentators and score their accuracy, aiming to reduce the "immunity to falsification" enjoyed by vague pundits.
- Prediction markets (like Robin Hanson's Futarchy) are viewed as promising tools, though Tetlock prefers individual forecasting tournaments to avoid the risk of market manipulation by well-funded actors.
Career & Future Directions:
- Ideal careers for aspiring forecasters involve a combination of social science, economics, and computer science/AI, often rooted in public policy schools.
- Tetlock emphasizes the need to bridge "rigor and relevance" in future research, moving from purely accuracy-focused questions to those that generate insightful, probative clusters on major global themes.
- Organizations like Good Judgment Incorporated, IARPA, and Open Philanthropy are identified as key hubs for work in this field.