newsfilter.io
Interview

#15 - Prof Tetlock on how chimps beat Berkeley undergrads and when it’s wise to defer to the wise

  • Research Foundation & Scoring:

    • Philip Tetlock's 35-year research program includes tens of thousands of participants and millions of predictions, primarily scored using the Brier score system developed by meteorologist Glenn Breyer.
    • Lower Brier scores indicate higher accuracy; a perfect score of 0 implies clairvoyance, while a score of 2 indicates "inverse clairvoyance" (perfectly wrong predictions).
    • A random-forecasting baseline (equivalent to a dart-throwing chimp or coin flip) yields a score of 0.5.
    • Initial findings revealed that political experts performed only marginally better than dart-throwing chimps, with experts failing to significantly outperform "attentive readers" of elite press (e.g., New York Times) and losing to simple extrapolation algorithms.
    • Experts significantly outperformed only one group: Berkeley undergraduates, who performed worse than chance due to extreme overconfidence and lack of self-awareness.
  • Key Behavioral Findings (Hedgehogs vs. Foxes):

    • "Hedgehogs" (those with one big theory) performed worse than "foxes" (those who stick to data and acknowledge multiple viewpoints), particularly in long-range forecasts.
    • Hedgehogs exhibited higher overconfidence, claiming events were 100% certain when they occurred 80% of the time, and events predicted at 80% occurred only 65% of the time.
    • Better forecasters displayed curiosity, open-mindedness, and a high tolerance for cognitive dissonance, allowing them to entertain counterfactuals that challenged their ideological priors.
    • Poor performers often relied on post-hoc defenses like "timing errors" or "exogenous shocks" rather than updating beliefs after failed predictions.
  • The "Superforecasters" & Aggregation:

    • The later Superforecasting research identified a subset of individuals who consistently made highly accurate predictions through rigorous training and teamwork.
    • Aggregating judgments from superforecasters produced results superior to those of intelligence community analysts with access to classified data.
    • Superforecasting teams utilized a weighted average of the best recent forecasts, "extremizing" the result based on the diversity of viewpoints among forecasters (cognitive diversity).
    • Extreme caution is advised when extremizing judgments within groups that have already collaborated extensively, as they risk "self-extremizing" rather than reflecting true cognitive diversity.
  • New Tournament (Hybrid Forecasting Competition):

    • A new IARPA-sponsored tournament pits humans, machines, and human-machine hybrids against each other to determine optimal forecasting methods.
    • The competition targets two domains: quantitative problems (where AI currently holds an advantage) and idiosyncratic, context-specific questions (e.g., Yasser Arafat's autopsy, Syrian civil war duration) where human intuition may still excel.
    • Volunteers can participate at hybridforecasting.com, where they are assigned to specific forecasting algorithms to test different aggregation strategies.
  • Calibration, Training, and Bias:

    • Calibration training shows limited transferability; skills learned in one domain (e.g., poker) do not automatically generalize to unrelated domains (e.g., geopolitics) unless the training includes understanding the underlying metrics of probability.
    • Tetlock's personal experience in a tournament demonstrated that "tweaking" optimal algorithmic predictions led to a drop from 2nd place to 35th, underscoring the value of deferring to established aggregation logic over individual cognitive effort.
    • Publicly committing to a strong position can freeze beliefs and trigger "defensive bolstering," making it harder to update views in the face of new evidence.
  • Forecasting Extreme Events ("The Tails"):

    • Assessing the accuracy of forecasts for low-probability, high-impact events (e.g., nuclear war, pandemics) is difficult due to the scarcity of actual events for validation.
    • Validity in the "tails" is often assessed through logical coherence checks, such as verifying "temporal scope sensitivity" (ensuring probabilities increase appropriately as the time horizon lengthens).
    • "Question clustering" is proposed as a method to predict long-term risks by aggregating answers to specific, diagnostic intermediate questions (e.g., driverless car adoption rates as a proxy for the Fourth Industrial Revolution).
  • Institutional & Societal Implications:

    • Tetlock opposes the "populist know-nothingism" that dismisses all expertise, arguing that ignoring experts is dangerous, though justified skepticism of overconfident pundits is necessary.
    • He distinguishes between domains where expertise should be highly deferred (e.g., surgery, where feedback is rapid and clear) versus macro-political domains where feedback is delayed and ambiguous.
    • The "Alpha Pundit Challenge" is a proposed initiative to extract implicit forecasts from media commentators and score their accuracy, aiming to reduce the "immunity to falsification" enjoyed by vague pundits.
    • Prediction markets (like Robin Hanson's Futarchy) are viewed as promising tools, though Tetlock prefers individual forecasting tournaments to avoid the risk of market manipulation by well-funded actors.
  • Career & Future Directions:

    • Ideal careers for aspiring forecasters involve a combination of social science, economics, and computer science/AI, often rooted in public policy schools.
    • Tetlock emphasizes the need to bridge "rigor and relevance" in future research, moving from purely accuracy-focused questions to those that generate insightful, probative clusters on major global themes.
    • Organizations like Good Judgment Incorporated, IARPA, and Open Philanthropy are identified as key hubs for work in this field.