Sampling
eBook - ePub

Sampling

Design and Analysis

  1. 654 pages
  2. English
  3. ePUB (mobile friendly)
  4. Available on iOS & Android
eBook - ePub

Sampling

Design and Analysis

About this book

"The level is appropriate for an upper-level undergraduate or graduate-level statistics major. Sampling: Design and Analysis (SDA) will also benefit a non-statistics major with a desire to understand the concepts of sampling from a finite population. A student with patience to delve into the rigor of survey statistics will gain even more from the content that SDA offers. The updates to SDA have potential to enrich traditional survey sampling classes at both the undergraduate and graduate levels. The new discussions of low response rates, non-probability surveys, and internet as a data collection mode hold particular value, as these statistical issues have become increasingly important in survey practice in recent years… I would eagerly adopt the new edition of SDA as the required textbook." (Emily Berg, Iowa State University)

What is the unemployment rate? What is the total area of land planted with soybeans? How many persons have antibodies to the virus causing COVID-19? Sampling: Design and Analysis, Third Edition shows you how to design and analyze surveys to answer these and other questions. This authoritative text, used as a standard reference by numerous survey organizations, teaches the principles of sampling with examples from social sciences, public opinion research, public health, business, agriculture, and ecology. Readers should be familiar with concepts from an introductory statistics class including probability and linear regression; optional sections contain statistical theory for readers familiar with mathematical statistics.

Key Features:

  • Has been thoroughly revised to incorporate recent research and applications.
  • Includes a new chapter on nonprobability samples, and more than 200 new examples and exercises have been added.
  • Teaches the principles of sampling with examples from social sciences, public opinion research, public health, business, agriculture, and ecology.

SDA's companion website contains data sets, computer code, and links to two free downloadable supplementary books (also available in paperback) that provide step-by-step guides—with code, annotated output, and helpful tips—for working through the SDA examples. Instructors can use either R or SAS® software.

  • SASÂŽ Software Companion for Sampling: Design and Analysis, Third Edition by Sharon L. Lohr (2022, CRC Press)
  • R Companion for Sampling: Design and Analysis, Third Edition by Yan Lu and Sharon L. Lohr (2022, CRC Press)

Information

Year
2021
Print ISBN
9780367279509
Edition
3
eBook ISBN
9781000478266

1Introduction

DOI: 10.1201/9780429298899-1
When statistics are not based on strictly accurate calculations, they mislead instead of guide. The mind easily lets itself be taken in by the false appearance of exactitude which statistics retain in their mistakes, and confidently adopts errors clothed in the form of mathematical truth.
—Alexis de Tocqueville, Democracy in America

1.1 Guidance from Samples

We all use data from samples to make decisions. When tasting soup to correct the seasoning, deciding to buy a book after reading the first page, choosing a major after taking first-year college classes, or buying a car following a test drive, we rely on partial information to judge the whole.
External data used to help with those decisions come from samples, too. Statistics such as the average rating for a book in online reviews, the median salary of psychology majors, the percentage of persons with an undergraduate mathematics degree who are working in a mathematics-related job, or the number of injuries resulting from automobile accidents in 2018 are all derived from samples. So are statistics about unemployment and poverty rates, inflation, number and characteristics of persons with diabetes, medical expenditures of persons aged 65 and over, persons experiencing food insecurity, criminal victimizations not reported to the police, reading proficiency among fourth-grade children, household expenditures on energy, public opinion of political candidates, land area under cultivation for rice, livestock owned by farmers, contaminants in drinking water, size of the Antarctic population of emperor penguins—I could go on, but you get the idea. Samples, and statistics calculated from samples, surround us.
But statistics from some samples are more trustworthy than those from others. What distinguishes, using Tocqueville's words beginning this page, statistics that “mislead” from those that “guide”?
This book sets out the statistical principles that tell you how to design a sample survey, and analyze data from a sample, so that statistics calculated from a sample accurately describe the population from which the sample was drawn. These principles also help you evaluate the quality of any statistic you encounter that originated from a sample survey.
Before embarking on our journey, let's look at how a statistic from a now-infamous survey misled readers in 1936.
Example 1.1. The Survey That Killed a Magazine. Any time a pollster predicts the wrong winner of an election, some commentator is sure to mention the Literary Digest Poll of 1936. It has been called “one of the worst political predictions in history” (Little, 2016) and is regularly cited as the classic example of poor survey practice. What went wrong with the poll, and was it really as flawed as it has been portrayed?
In the first three decades of the twentieth century, The Literary Digest, a weekly news magazine founded in 1890, was one of the most respected news sources in the United States. In presidential election years, it, like many other newspapers and magazines, devoted page after page to speculation about who would win the election. For the 1916 election, however, the editors wrote that “[p]olitical forecasters are in the dark” and asked subscribers in five states to mail in a ballot indicating their preferred candidate (Literary Digest, 1916).
The 1916 poll predicted the correct winner in four of the five states, and the magazine continued polling subsequent presidential elections, with a larger sample each time. In each of the next four election years—1920, 1924 (the first year the poll collected data from all states), 1928, and 1932—the person predicted to win the presidency did so, and the magazine accurately predicted the margin of victory. In 1932, for example, the poll predicted that Franklin Roosevelt would receive 56% of the popular vote and 474 votes in the Electoral College; in the actual election, Roosevelt received 57% of the popular vote and 472 votes in the Electoral College.
With such a strong record of accuracy, it is not surprising that the editors of The Literary Digest gained confidence in their polling methods. Launching the 1936 poll, they wrote:
The Poll represents thirty years' constant evolution and perfection. Based on the “commercial sampling” methods used for more than a century by publishing houses to push book sales, the present mailing list is drawn from every telephone book in the United States, from the rosters of clubs and associations, from city directories, lists of registered voters, classified mail-order and occupational data. (Literary Digest, 1936b, p. 3)
On October 31, 1936, the poll predicted that Republican Alf Landon would receive 54% of the popular vote, compared with 41% for Democrat Franklin Roosevelt. The final article on polling before the election contained the statement, “We make no claim to infallibility. We did not coin the phrase ‘uncanny accuracy’ which has been so freely applied to our Polls” (Literary Digest, 1936a). It is a good thing The Literary Digest made no claim to infallibility. In the election, Roosevelt received 61% of the vote; Landon, 37%. It is widely thought that this polling debacle contributed to the demise of the magazine in 1938.
What went wrong? One problem may have been that names of persons to be polled were compiled from sources such as telephone directories and automobile registration lists. Households with a telephone or automobile in 1936 were generally more affluent than other households, and opinion of Roosevelt's economic policies was generally related to the economic class of the respondent. But the mailing list's deficiencies do not explain all of the difference. Postmortem analyses of the poll (Squire, 1988; Calahan, 1989; Lusinchi, 2012) indicated that even persons with both a car and a telephone tended to favor Roosevelt, though not to the degree that persons with neither car nor telephone supported him.
Nonresponse—the failure of persons selected for the sample to provide data—was likely the source of much of the error. Ten million questionnaires were mailed out, and more than 2.3 million were returned—an enormous sample, but fewer than one-quarter of those solicited. In Allentown, Pennsylvania, for example, the survey was mailed to every registered voter, but the poll results for Allentown were still incorrect because only one-third of the ballots were returned (Literary Digest, 1936c). Squire (1988) reported that persons supporting Landon were much more likely to have returned the survey; in fact, many Roosevelt supporters did not remember receiving a survey even though they were on the mailing list.
One lesson to be learned from The Literary Digest poll is that the sheer size of a sample is no guarantee of its accuracy. The Digest editors became complacent because they sent out questionnaires to more than one-quarter of all registered voters and obtained a huge sample of more than 2.3 million people. But large unrepresentative samples can perform as badly as small unrepresentative samples. A large unrepresentative sample may even do more harm than a small one because many people think that large samples are always superior to small ones. In reality, as we shall discuss in this book, the design of the sample survey—how units are selected to be in the sample—is far more important than its size.
Another lesson is that past accuracy of a flawed sampling procedure does not guarantee future results. The Literary Digest poll was accurate for five successive elections—until suddenly, in 1936, it wasn't. Reliable statistics result from using statistically sound sampling and estimation procedures. With good procedures, statisticians can provide a measure of a statistic's accuracy; without good procedures, a sampling disaster can happen at any time even if previous statistics appeared to be accurate. ■
Some of today's data sets make the size of the Literary Digest's sample seem tiny by comparison, and some types of data can be gathered almost instantaneously from all over the world. But the challenges of inferring the characteristics of a population when we observe only part of it remain the same. The statistical principles underlying sampling apply to any sample, of any size, at any time or place in the universe.
Chapters 2 through 7 of this book show you how to design a sample so that its data can be used to estimate characteristics of unobserved parts of the population; Chapters 9 through 14 show how to use survey data to estimate population sizes, relationships among variables, and other characteristics of interest. But even though you might design and select your sample in accordance with statistical principles, in many cases, you cannot guarantee that everyone selected for the sample will agree to participate in it. A typical election poll in 2021 has a much lower response rate than the Literary Digest poll, but modern survey samplers use statistical models, described in Chapters 8 and 15, to adjust for the nonresponse. We'll return to the Literary Digest poll in Chapter 15 and see if a nonresponse model would have improved the poll's forecast (and perhaps have saved the magazine).

1.2 Populations and Representative Samples

In the 1947 movie “Magic Town,” the public opinion researcher played by James Stewart discovered a town that had exactly the same characteristics as the whole United States: Grandview had exactly the same proportion of people who voted Republican, the same proportion of people under the poverty line, the same proportion of auto mechanics, and so on, as the United States taken as a whole. All that Stewart's character had to do was to interview the people of Grandview, and he would know public opinion in the United States.
Grandview is a “scaled-down” version of the population, mirroring every characteristic of the whole population. In that sense, it is representative of the population of the United States because any numerical quantity that could be calculated from the population can be inferred from the sample.
But a sample does not necessarily have to be a small-scale replica of the population to be representative. As we shall discuss in Chapters 2 and 3, a sample is representative if it can be used to “reconstruct” what the population looks like—and if we can provide an accurate assessment of how good that reconstruction is.
Some definitions are needed to make the notions of a “population” and a “representative sample” more precise.
  • Observation unit An object on which a measurement is taken, sometimes called an element. In surveys of human populations, observation units are often individual persons; in agriculture or ecology surveys, they may be small areas of land; in audit surveys, they may be financial records.
  • Target population The complete collection of observations we want to study. Defining the target population is an important and often d...

Table of contents

  1. Cover Page
  2. Half-Title Page
  3. Series Page
  4. Title Page
  5. Copyright Page
  6. Dedication Page
  7. Contents
  8. Preface
  9. Symbols and Acronyms
  10. 1 Introduction
  11. 2 Simple Probability Samples
  12. 3 Stratified Sampling
  13. 4 Ratio and Regression Estimation
  14. 5 Cluster Sampling with Equal Probabilities
  15. 6 Sampling with Unequal Probabilities
  16. 7 Complex Surveys
  17. 8 Nonresponse
  18. 9 Variance Estimation in Complex Surveys
  19. 10 Categorical Data Analysis in Complex Surveys
  20. 11 Regression with Complex Survey Data
  21. 12 Two-Phase Sampling
  22. 13 Estimating the Size of a Population
  23. 14 Rare Populations and Small Area Estimation
  24. 15 Nonprobability Samples
  25. 16 Survey Quality
  26. A Probability Concepts Used in Sampling
  27. Bibliography
  28. Index

Trusted by 375,005 students

Access to over 1.5 million titles for a fair monthly price.

Study more efficiently using our study tools.

Frequently asked questions

Yes, you can cancel anytime from the Subscription tab in your account settings on the Perlego website. Your subscription will stay active until the end of your current billing period. Learn how to cancel your subscription
No, books cannot be downloaded as external files, such as PDFs, for use outside of Perlego. However, you can download books within the Perlego app for offline reading on mobile or tablet. Learn how to download books offline
Perlego offers two plans: Essential and Complete
  • Essential is ideal for learners and professionals who enjoy exploring a wide range of subjects. Access the Essential Library with 800,000+ trusted titles and best-sellers across business, personal growth, and the humanities. Includes unlimited reading time and Standard Read Aloud voice.
  • Complete: Perfect for advanced learners and researchers needing full, unrestricted access. Unlock 1.5M+ books across hundreds of subjects, including academic and specialized titles. The Complete Plan also includes advanced features like Premium Read Aloud and Research Assistant.
Both plans are available with monthly, semester, or annual billing cycles.
We are an online textbook subscription service, where you can get access to an entire online library for less than the price of a single book per month. With over 1.5 million books across 990+ topics, we’ve got you covered! Learn about our mission
Look out for the read-aloud symbol on your next book to see if you can listen to it. The read-aloud tool reads text aloud for you, highlighting the text as it is being read. You can pause it, speed it up and slow it down. Learn more about Read Aloud
Yes! You can use the Perlego app on both iOS and Android devices to read anytime, anywhere — even offline. Perfect for commutes or when you’re on the go.
Please note we cannot support devices running on iOS 13 and Android 7 or earlier. Learn more about using the app
Yes, you can access Sampling by Sharon L. Lohr in PDF and/or ePUB format, as well as other popular books in Social Sciences & Probability & Statistics. We have over 1.5 million books available in our catalogue for you to explore.