Interview
#15 - Prof Tetlock on how chimps beat Berkeley undergrads and when it’s wise to defer to the wise
- IARPA is actively running the Hybrid Forecasting Competition, a crowdsourcing tournament pitting humans, machines, and hybrids against one another, with volunteer recruitment ongoing at hybridforecasting.com.
- Machines are projected to hold an advantage in complex quantitative forecasting, such as OECD economic growth patterns, but may struggle with idiosyncratic, context-specific events like the Syrian Civil War duration where statistical base rates are elusive; human forecasters are expected to compete by detecting turbulence and adjusting for model overfitting.
- Statistical algorithms utilizing weighted averages of top forecasters with extremization based on cognitive diversity are expected to remain the winning approach, though applying further extremization to superforecasting teams may yield inferior results compared to their self-extremized collaborative judgments.
- Matching the accuracy of a single superforecaster is estimated to require between 10 and 35 regular forecasters, depending on the specific question and group composition.
- Calibration training is expected to yield modest transfer effects across unrelated domains, such as moving from poker to weather or interest rates, with greater transfer anticipated only if forecasters deeply understand calibration and resolution metrics.
- Strategic interventions in algorithms can degrade performance; reliance on a top algorithm without personal tweaking in a prior tournament was estimated to have secured a top-five ranking, whereas cognitive intervention resulted in approximately 35th place.
- The "inside view" in forecasting presents risks of cognitive conservatism or excess volatility, yet a categorical prohibition is considered too extreme for practical application.
- Nate Silver's 70% probability estimate for the 2016 Hillary Clinton victory is expected to be viewed as a prudent judgment that avoided extremizing errors derived from correlated poll measurement errors.
- Making strong public political commitments is expected to freeze attitudes, creating cognitive and emotional barriers to updating beliefs in response to new evidence due to defensive bolstering.
- Feeding accurate probabilities to decision-makers with reckless utility functions may lead to negative outcomes if those actors act risk-seeking based on the intelligence provided.
- Assessing forecast accuracy for tail events (less than 5% likelihood) is expected to be extremely difficult due to the scarcity of observed events required to generate a meaningful track record.
- Low-probability forecasts are expected to be evaluated for logical coherence by checking for temporal scope insensitivity, where forecasters incorrectly assign similar probabilities to events occurring over different timeframes, such as six months versus four years.
- Future forecasting tournaments are expected to bridge the rigor and relevance divide by developing question clusters that are rigorously measurable yet diagnostic of broader societal themes like U.S.-China relations.
- The creation of a dedicated high school or university course to train superforecasters is expected to be a feasible project potentially fundable by institutions like Minerva.
- The financial industry is expected to increase interest in forecasting research as the field progresses beyond the initial discovery of teachable probabilistic accuracy techniques.
- The Alpha Pundit Challenge is expected to eventually cause pundits to become more circumspect in their forecasts due to fears regarding credibility tracking against imputed probability ranges.
- Current political polarization and politicization are expected to subside over the long arc of history as the competitive advantages of intellectual temperance and accurate forecasting become recognized.
- Institutions guided by realistic probability estimates are expected to outperform those that do not rely on accurate forecasting over the long term.