newsfilter.io
Webinar

Election polling: why is it so difficult?

  • Historical Evolution of Polling Accuracy

    • Pre-1936 polling relied on "straw polls" sampling local crowds (e.g., holiday celebrations, militia meetings), which were non-representative and effectively guesswork.
    • The 1936 Literary Digest failure, which predicted Alf Landon's victory over FDR using skewed wealth-based samples, proved the necessity of random sampling.
    • George Gallup's 1936 success, using a smaller but demographically representative random sample, established the modern scientific polling standard.
    • Election betting markets were historically significant in the 1800s and 1900s before newspapers adopted polling methods.
  • Modern Methodological Standards

    • Scientific polling requires hitting specific demographic quotas (age, gender, education, social class) to mirror the total population.
    • Question phrasing remains consistent globally, typically asking, "If there were an election tomorrow, how would you vote?"
    • Pollsters apply weighting adjustments to correct for over-representation of politically engaged respondents, particularly regarding education levels.
  • Analysis of the 2016 U.S. Election Failure

    • National polls in 2016 were accurate within the historical margin of error (1.2-point error: Clinton won 51.1% vs. predicted 52.3%).
    • The election loss stemmed from state-level polling failures in Michigan, Wisconsin, and Pennsylvania, where polls were up to 10 points off.
    • The inquiry identified a lack of non-college graduate respondents in Midwest state polls, skewing results away from the Trump coalition.
    • Late-deciding voters in the final week of the election were not captured by pollsters, contributing to the discrepancy.
    • The 1948 U.S. election serves as a historical precedent for polling failure, with major outlets declaring Dewey the winner before Truman's victory.
  • Statistical Modeling and Forecasting Approaches

    • Forecasting models utilize two primary inputs: direct polling data and "fundamentals" (economic growth, unemployment, candidate funding).
    • Fundamental-based models can accurately predict outcomes (e.g., 1992 U.S. election) even when individual polls are inconsistent.
    • The 2022 French presidential election was modeled using a chained, step-by-step approach to account for its two-round voting system.
    • The methodology employed a "spline" statistical tool to average daily polls and historical margin-of-error data (e.g., polls are typically 4 points off 100 days prior).
    • 10 million simulations were run to generate probability distributions for election outcomes rather than binary predictions.
  • 2022 French Election Prediction Outcomes

    • On February 5, 2022 (10 weeks prior), simulations predicted a 92% probability of Emmanuel Macron reaching the runoff and a 48% probability for Marine Le Pen.
    • In simulated head-to-head matchups, Macron won 88% of the time.
    • Two months before the election, Macron's victory probability stood at approximately 80%.
  • Forward-Looking Statements and Limitations

    • Models are theoretically incapable of achieving 100% accuracy; they aim to use data to predict outcomes "pretty close" to the actual result.
    • Forecast reliability is directly constrained by the quality of the underlying polls.
    • Extreme external events (e.g., pandemics, wars) introduce variables that models cannot fully predict.
    • As historical data accumulates, election modeling precision is expected to improve.
    • The source code for the models and detailed methodology are publicly available via provided links.