Panel, Conference Presentation
The Art and Science of Polling
Milken InstituteMollyann Brodie, Jon Cohen, Michael Dimock, Alex Lundry, Margie Omero, Molly Ambroti
- Public engagement and horror regarding the election are expected to persist throughout the cycle, with clear nominee identification anticipated by election day due to Clinton's insurmountable delegate lead and Trump's projected strength in states like Indiana.
- The Republican primary is predicted to remain highly contentious with divisive issues brought to the foreground, while the Democratic side is expected to focus on issues favored within the party, though the impact of candidates dropping out on divisiveness remains uncertain before and after the California primary.
- Polling accuracy for primaries is predicted to improve when elections resemble previous cycles and group support clarifies, whereas predictive models face significant failure risks when conditions change, as evidenced by past misses in Michigan and Iowa despite generally strong industry performance.
- Data science has transformed the sector through machine learning and deep learning, driving a shift from pure prediction (estimated at 5% of tasks) to strategy and tactics, with viability influencing voter support and "horse race" polling potentially creating self-fulfilling narratives.
- The industry is experiencing a 250 percent increase in online Republican primary polls compared to the last cycle, while telephone response rates are declining; high-quality surveys are becoming a luxury item achievable only at extraordinary cost, prompting a trend toward aggregating multiple polls into summary numbers.
- Significant polling shifts are noted, including an 11-point rise for Trump in the national primary over two weeks driven by perceived winning momentum, though his ultimate ceiling remains unknown, and a strong correlation of 0.9 plus exists between Trump's news coverage and polling numbers.
- Methodological challenges include the difficulty of using narrow margins of error (±3% or 4%) for debate placement, the risk of "garbage in, garbage out" when mixing data quality levels, and the breakdown of consensus on survey quality, though online metrics like Google search volume and social media share of voice provided early indicators of Trump's viability before polls caught up in October.
- Qualitative research and likely voter modeling are shifting toward unstructured web data and machine-driven historical turnout models rather than self-reports, while the immediate media landscape encourages over-reading small poll changes and bans on pre-election polling are deemed unlikely due to constitutional and stakeholder interests.
- Specific campaign failures, such as the Jeb campaign's data mining which failed to predict the Trump "black swan" event due to reliance on prior experience, contrast with public polling data that eventually confirmed voters intended to support Trump on core issues like immigration and terrorism.
- Bias and misinterpretation risks persist where analysts may insert their own policy assumptions into questions, though aggregating known biases in the data explosion is viewed as beneficial for models, and the essential democratic role of polls remains understanding the "why" behind voter behavior despite the noise of immediate media cycles.