We all use data from samples to make decisions. When tasting soup to correct the seasoning, deciding to buy a book after reading the first page, choosing a major after taking first-year college classes, or buying a car following a test drive, we rely on partial information to judge the whole.
External data used to help with those decisions come from samples, too. Statistics such as the average rating for a book in online reviews, the median salary of psychology majors, the percentage of persons with an undergraduate mathematics degree who are working in a mathematics-related job, or the number of injuries resulting from automobile accidents in 2018 are all derived from samples. So are statistics about unemployment and poverty rates, inflation, number and characteristics of persons with diabetes, medical expenditures of persons aged 65 and over, persons experiencing food insecurity, criminal victimizations not reported to the police, reading proficiency among fourth-grade children, household expenditures on energy, public opinion of political candidates, land area under cultivation for rice, livestock owned by farmers, contaminants in drinking water, size of the Antarctic population of emperor penguinsâI could go on, but you get the idea. Samples, and statistics calculated from samples, surround us.
But statistics from some samples are more trustworthy than those from others. What distinguishes, using Tocqueville's words beginning this page, statistics that âmisleadâ from those that âguideâ?
This book sets out the statistical principles that tell you how to design a sample survey, and analyze data from a sample, so that statistics calculated from a sample accurately describe the population from which the sample was drawn. These principles also help you evaluate the quality of any statistic you encounter that originated from a sample survey.
Before embarking on our journey, let's look at how a statistic from a now-infamous survey misled readers in 1936.
Example 1.1. The Survey That Killed a Magazine. Any time a pollster predicts the wrong winner of an election, some commentator is sure to mention the Literary Digest Poll of 1936. It has been called âone of the worst political predictions in historyâ (Little, 2016) and is regularly cited as the classic example of poor survey practice. What went wrong with the poll, and was it really as flawed as it has been portrayed?
In the first three decades of the twentieth century, The Literary Digest, a weekly news magazine founded in 1890, was one of the most respected news sources in the United States. In presidential election years, it, like many other newspapers and magazines, devoted page after page to speculation about who would win the election. For the 1916 election, however, the editors wrote that â[p]olitical forecasters are in the darkâ and asked subscribers in five states to mail in a ballot indicating their preferred candidate (Literary Digest, 1916).
The 1916 poll predicted the correct winner in four of the five states, and the magazine continued polling subsequent presidential elections, with a larger sample each time. In each of the next four election yearsâ1920, 1924 (the first year the poll collected data from all states), 1928, and 1932âthe person predicted to win the presidency did so, and the magazine accurately predicted the margin of victory. In 1932, for example, the poll predicted that Franklin Roosevelt would receive 56% of the popular vote and 474 votes in the Electoral College; in the actual election, Roosevelt received 57% of the popular vote and 472 votes in the Electoral College.
With such a strong record of accuracy, it is not surprising that the editors of The Literary Digest gained confidence in their polling methods. Launching the 1936 poll, they wrote:
The Poll represents thirty years' constant evolution and perfection. Based on the âcommercial samplingâ methods used for more than a century by publishing houses to push book sales, the present mailing list is drawn from every telephone book in the United States, from the rosters of clubs and associations, from city directories, lists of registered voters, classified mail-order and occupational data. (Literary Digest, 1936b, p. 3)
On October 31, 1936, the poll predicted that Republican Alf Landon would receive 54% of the popular vote, compared with 41% for Democrat Franklin Roosevelt. The final article on polling before the election contained the statement, âWe make no claim to infallibility. We did not coin the phrase âuncanny accuracyâ which has been so freely applied to our Pollsâ (Literary Digest, 1936a). It is a good thing The Literary Digest made no claim to infallibility. In the election, Roosevelt received 61% of the vote; Landon, 37%. It is widely thought that this polling debacle contributed to the demise of the magazine in 1938.
What went wrong? One problem may have been that names of persons to be polled were compiled from sources such as telephone directories and automobile registration lists. Households with a telephone or automobile in 1936 were generally more affluent than other households, and opinion of Roosevelt's economic policies was generally related to the economic class of the respondent. But the mailing list's deficiencies do not explain all of the difference. Postmortem analyses of the poll (Squire, 1988; Calahan, 1989; Lusinchi, 2012) indicated that even persons with both a car and a telephone tended to favor Roosevelt, though not to the degree that persons with neither car nor telephone supported him.
Nonresponseâthe failure of persons selected for the sample to provide dataâwas likely the source of much of the error. Ten million questionnaires were mailed out, and more than 2.3 million were returnedâan enormous sample, but fewer than one-quarter of those solicited. In Allentown, Pennsylvania, for example, the survey was mailed to every registered voter, but the poll results for Allentown were still incorrect because only one-third of the ballots were returned (Literary Digest, 1936c). Squire (1988) reported that persons supporting Landon were much more likely to have returned the survey; in fact, many Roosevelt supporters did not remember receiving a survey even though they were on the mailing list.
One lesson to be learned from The Literary Digest poll is that the sheer size of a sample is no guarantee of its accuracy. The Digest editors became complacent because they sent out questionnaires to more than one-quarter of all registered voters and obtained a huge sample of more than 2.3 million people. But large unrepresentative samples can perform as badly as small unrepresentative samples. A large unrepresentative sample may even do more harm than a small one because many people think that large samples are always superior to small ones. In reality, as we shall discuss in this book, the design of the sample surveyâhow units are selected to be in the sampleâis far more important than its size.
Another lesson is that past accuracy of a flawed sampling procedure does not guarantee future results. The Literary Digest poll was accurate for five successive electionsâuntil suddenly, in 1936, it wasn't. Reliable statistics result from using statistically sound sampling and estimation procedures. With good procedures, statisticians can provide a measure of a statistic's accuracy; without good procedures, a sampling disaster can happen at any time even if previous statistics appeared to be accurate. â
Some of today's data sets make the size of the Literary Digest's sample seem tiny by comparison, and some types of data can be gathered almost instantaneously from all over the world. But the challenges of inferring the characteristics of a population when we observe only part of it remain the same. The statistical principles underlying sampling apply to any sample, of any size, at any time or place in the universe.
Chapters 2 through 7 of this book show you how to design a sample so that its data can be used to estimate characteristics of unobserved parts of the population; Chapters 9 through 14 show how to use survey data to estimate population sizes, relationships among variables, and other characteristics of interest. But even though you might design and select your sample in accordance with statistical principles, in many cases, you cannot guarantee that everyone selected for the sample will agree to participate in it. A typical election poll in 2021 has a much lower response rate than the Literary Digest poll, but modern survey samplers use statistical models, described in Chapters 8 and 15, to adjust for the nonresponse. We'll return to the Literary Digest poll in Chapter 15 and see if a nonresponse model would have improved the poll's forecast (and perhaps have saved the magazine).