Pew Research Center

FOR RELEASE MAY 18, 2001

Screening Likely Voters: A Survey Experiment

Introduction and Summary Traditionally, pollsters trying to accurately assess voter intentions have struggled with a basic problem — figuring out who actually is going to show up to vote. In the 2000 election campaign, sharp fluctuations in the Gallup Organization’s daily tracking poll were blamed by some on difficulties in nailing down likely voters. Similar […]

FOR MEDIA OR OTHER INQUIRIES:

Communications Department
202.419.4372
www.pewresearch.org

RECOMMENDED CITATION

Pew Research Center, May 2001, "Screening Likely Voters: A Survey Experiment"

About Pew Research Center

Pew Research Center is a nonpartisan, nonadvocacy fact tank that informs the public about the issues, attitudes and trends shaping the world. It does not take policy positions. The Center conducts public opinion polling, demographic research, computational social science research and other data-driven research. It studies politics and policy; news habits and media; the internet and technology; religion; race and ethnicity; international affairs; social, demographic and economic trends; science; research methodology and data science; and immigration and migration. Pew Research Center is a subsidiary of The Pew Charitable Trusts, its primary funder.

© Pew Research Center 2026

Table of contents

  • About Pew Research Center
  • Screening Likely Voters: A Survey Experiment
  • Methodology

Screening Likely Voters: A Survey Experiment

Introduction and Summary Traditionally, pollsters trying to accurately assess voter intentions have struggled with a basic problem — figuring out who actually is going to show up to vote. In the 2000 election campaign, sharp fluctuations in the Gallup Organization’s daily tracking poll were blamed by some on difficulties in nailing down likely voters. Similar […]

Introduction and Summary

Traditionally, pollsters trying to accurately assess voter intentions have struggled with a basic problem — figuring out who actually is going to show up to vote. In the 2000 election campaign, sharp fluctuations in the Gallup Organization’s daily tracking poll were blamed by some on difficulties in nailing down likely voters. Similar complaints arose during the 1998 congressional elections, when some critics of President Clinton charged that likely voter samples included too many Democrats sympathetic to Clinton. In 1999, the Pew Research Center undertook an experiment to study the accuracy of the likely voter models employed for decades by leading survey organizations, including Gallup and the Pew Research Center. In this comprehensive experiment, we set out to discover how many of the voters we classified as “likely” actually voted. The Center used as a test case the closely-contested 1999 mayoral race in Philadelphia. It was one of that city’s closest ever — just 9,447 votes separated the victor, Democrat John Street, from his Republican rival Sam Katz. Polling was conducted in two waves among 2,415 registered voters: one survey was taken two weeks before the election, the second was conducted in the last week before voters went to the polls. Aside from the usual battery of questions assessing voting preferences and intentions of voting, participants were asked to provide their name and address — information which was used to match pre-election survey responses with actual voting records. Overall, we were able to match 70% of registered voters polled with Philadelphia voting records to determine whether they actually voted. What we discovered was that the traditional methods originally developed by Gallup in the 1950s to sort voters from non-voters still work reasonably well, particularly when compared to the alternatives. This method uses an eight-item likely voter index designed to assess not only a voter’s preferences, but their past history of voting, interest in the campaign and knowledge of where to vote on Election Day. Using this index, the Center correctly predicted the voting behavior of 73% of registered voters.

Clearly, the index does not forecast the behavior of all respondents. The 73% accuracy rate means that 27% of respondents were wrongly classified — those who were determined as unlikely to vote but cast ballots (17%), or non-voters who were misclassified as likely to vote (10%). But the likely voter index successfully identified the preferences of respondents who actually went to the polls, according to the validation study. In the second wave of polling, conducted the week before Election Day, Street held a three-point lead among registered voters. But the dead heat among likely voters more accurately reflected the split among those respondents who actually voted. More important, the result virtually mirrors the findings of a similar voter validation study conducted by Gallup during the 1984 presidential election, which correctly classified 69% of registered voters. While the polling business has undergone massive changes since then, the likely voter index remains a model of consistency. Here is a summary of our principal findings. A more comprehensive analysis by Pew Center survey director Michael Dimock is the subject of a paper presented May 19 at the annual conference of the American Association for Public Opinion Research. (Complete paper; You can also contact Michael Dimock via email at mdimock@pewresearch.org).

One of the main successes of the likely-voter index is in identifying a pool of respondents which, if not a perfect replica of the electorate, shares a similar demographic profile with the voting public. It gives a better approximation of the electorate than using registered voters. The survey of likely voters conducted two weeks before the election was nearly as accurate in assessing voters’ preferences as the one taken just prior to the election. Still, it is possible that the effectiveness of the likely voter screen would decline if used too early in an election campaign. In 1998, the Democratic composition of the likely voter poll changed significantly between September and Election Day. Using a single measure to determine likely voters increases the risk of bias and subjects estimates to the vagaries of individual elections. Some individual questions, such as those relating to actual voting intention (“Do you plan to vote?”) result in too many respondents being classified as likely voters, thus increasing the chance of a Democratic bias. Other rifle-shot questions — focusing on a respondent’s level of attention to the campaign, for example, — may classify too few people as likely voters, resulting in a Republican bias. Interestingly, the Pew experiment showed that, in particular, an individual question related to a candidate’s strength of support (how strongly the respondent supported the candidate) is not closely associated with voter turnout. Registered voters who expressed no preference in pre-election surveys were nearly as likely to vote as those who voiced strong support for candidates (and those with no preference were more likely to vote than those expressing moderate candidate support).

Clearly, Gallup’s index is not the only way of identifying likely voters — other indices, employing as few as four or as many as 15 questions, also are effective. These indices are superior to relying solely on samples of registered voters, and also are more effective than using regression analysis to determine a probability of voting for each respondent.

Methodology

The data used in this analysis are from two telephone surveys (Wave 1, Oct 13-21, 1999; Wave 2, Oct 27-30, 1999) conducted in the City of Philadelphia by the Pew Research Center for the People & the Press, under the direction of Schulman, Ronca and Bucuvalas, Inc. Each wave of the survey consists of approximately 1,600 interviews, drawn from two distinct samples (see below for details). Roughly two-thirds of respondents in each wave were drawn from a standard random-digit sample of telephone numbers selected from telephone exchanges in the City of Philadelphia. The random digit aspect of the sample is used to avoid “listing” bias and provide representation of both listed and unlisted numbers (including not-yet-listed). The design of the sample ensures this representation by random generation of the last two digits of telephone numbers selected on the basis of their telephone exchange and bank number. The other third of each wave is drawn directly from voter registration lists maintained by local government agencies. Registration lists were used to identify households encompassing at least one registered voter, with standard household randomization applied once telephone contact was made. This alternative sampling methodology was utilized to test whether “matching” survey respondents to voter registration lists more or less efficient using different sampling techniques. Though there are many possible sources of bias in the voter-list sample (not all registered voters provide a phone number when registering, registration records may not be completely up-to-date), the respondents drawn from each separate sampling procedure were similar in most demographic and political characteristics. For both RDD and listed samples, numbers were released for interviewing in replicates. Using replicates to control the release of sample to the field ensures that the complete call procedures are followed for the entire sample. At least 10 attempts were made to complete an interview at every sampled telephone number. The calls were staggered over times of day and days of the week to maximize the chances of making a contact with a potential respondent. All interview breakoffs and refusals were recontacted at least once in order to attempt to convert them to completed interviews. In each contacted household, interviewers asked to speak with the “youngest male 18 or older who is at home.” If there is no eligible man at home, interviewers asked to speak with “the oldest woman 18 or older who is at home.” This systematic respondent selection technique has been shown empirically to produce samples that closely mirror the population in terms of age and gender. Survey respondents were matched to voter registration lists after election day to validate their voting behavior. The matching process took into account five parameters: phone number, first name, last name, address, and the respondent’s age. Overall, we successfully matched 70% of respondents who identified themselves as registered voters to the voter registration lists (68% from the combined RDD samples, 75% from the combined listed samples). The inability to match 30% of respondents who claim to be registered reflects three factors, each of which might affect the representativeness of the sample. First, many respondents overreport voter registration. Second, many respondents refused to give their name and address, making matching difficult or impossible. Third, registration lists maintained by local government agencies may not be completely up-to-date. Non-response in telephone interview surveys produces some known biases in survey-derived estimates because participation tends to vary for different subgroups of the population. Both respondents and non-responding households were “matched” to voter registration lists, in order to gauge the relationship between survey participation and turnout. Though each wave of telephone interviewing is drawn from two separate sampling frames, the analysis of likely voter methodology is based only on matched cases which, in effect, are all drawn from the same sampling frame of the registration lists of local government agencies. As a result, all analysis is conducted on the combined listed and RDD samples. Data are not weighted to census parameters due to the fact that the registration lists do not represent a random distribution of the city’s population.