Patricio Ezdebski
Research Mentor(s): Josh Pasek
Mentor Department: Center for Political Studies
Authors: Patricio Ezdebski, Josh Pasek
Session: Session 3 (11:00am – 11:50am)
Presentation Type: Poster 90
Abstract
After the 2016 U.S. presidential election, many Americans’ faith in polling wavered. Polling data pointed to a decisive Clinton win in battleground states (particularly in the Rust Belt)—the exact opposite of what actually happened. Pre-election polling suffered another hit in 2020, when Biden eked out a win, despite being up by more than the margin of error in most swing states. The purpose of this research project is to evaluate the difference between the actual election results and the polling data by analyzing the mode and voter likelihood models used in polls and seeing if there was any correlation to the polls’ accuracy. To accomplish this task, election polling data was scraped from three main sources: RealClearPolitics, 538, and Wikipedia (for historical polling data). AI tools were used to download information about the polls from news stories and reports into Dropbox, where AI tools were used again to code the data. The two main variables that this poster is analyzing are the mode (the data collection method) and the voter likelihood model (how the poll identified individuals likely/not likely to vote). After the polls were downloaded, they were put on a spreadsheet, where the variables were manually coded by humans. Comparing the AI coding with the human codingreveals how AI tools can help to identify key features of polling but also where some prove insufficient. The implications of these findings are tremendously impactful for the uncertain future of the polling industry, which has needed to rebuild trust and confidence amongst the general public after multiple misses, and for AI companies, such as OpenAI, to further improve their models. These findings will hopefully be utilized to ensure more accurate polling in future U.S. presidential elections (and others).



