Sampling & Bias

Research Methodology

Chapter 7 · Sampling & Bias

A larger sample isn't automatically a better one. This chapter covers real sampling methods, the most famous real polling failure in the history of statistics — and an honest reanalysis decades later showing even that famous story was more complicated than it's usually told.

Probability vs. Non-Probability Sampling

Probability samplingNon-probability sampling
Every member of the population has a known chance of selectionSelection isn't random — some members have no real chance of being included
Simple random, stratified, cluster, systematicConvenience, quota, snowball
Genuinely supports statistical generalization to the full populationFaster and cheaper, but real generalization is genuinely limited

A Real, Canonical Case of Selection Bias

1936: The Literary Digest's Real Prediction

The magazine Literary Digest mailed roughly 10 million ballots — drawn from telephone directories, car registration lists, and its own subscriber rolls — and received about 2.3 million responses. It confidently predicted Alf Landon would beat Franklin Roosevelt, 57% to 43%. Roosevelt actually won, 62% to 38%.

Phones and cars were luxury items during the Great Depression, so the sample frame itself skewed toward wealthier respondents — a group genuinely less likely to support Roosevelt's New Deal policies than Landon's platform. George Gallup, using a real, much smaller but properly constructed sample, correctly predicted a Roosevelt win — a result now treated as the founding moment of modern scientific polling.

An Honest Complication — The Popular Story Isn't the Whole Story
A real 1988 reanalysis by political scientist Peverill Squire found that sample skew alone doesn't fully explain the error — even accounting for who was mailed a ballot, the poll would still have predicted Roosevelt correctly if everyone who received one had actually responded. The real, larger factor was non-response bias: Landon supporters returned their ballots at a meaningfully higher rate than Roosevelt supporters. The famous "the sample frame was rigged" story is real, but the size of its actual effect is smaller than commonly told — non-response was the bigger real driver.

Confirmation Bias

Psychologist Peter Wason's 1960 study introduced the "2-4-6 task": participants were told a rule governed which number triples were valid, shown one valid example, and asked to discover the rule by proposing their own triples. Most participants formed an overly narrow hypothesis and then only tested triples that would confirm it, rarely trying an example designed to prove themselves wrong. Wason coined the term confirmation bias for this real, well-documented tendency to favor evidence that supports an existing belief over evidence that challenges it.

The Real, Practical Lesson
Sampling bias and confirmation bias are two different failure points that reinforce each other. A biased sample can produce a confidently wrong result on its own — but a researcher who's already confident in a conclusion is also less likely to go looking for the kind of disconfirming case (a Gallup-style smaller, representative sample; a genuinely skeptical reanalysis like Squire's) that would catch the error. Both require the same real discipline: actively seeking out the evidence that could prove you wrong, not just the evidence that agrees with you.

Hands-On Exercises

Exercise 1

A researcher wants to survey "public opinion on a new city park proposal" by standing outside the park itself and asking passersby. Identify what kind of sampling this is, and explain, in your own words, why the resulting sample likely can't represent the opinions of city residents who never visit the park.

📄 View solution
Exercise 2

Explain, in your own words, why Squire's 1988 reanalysis is a genuinely important addition to the Literary Digest story, rather than just a minor footnote — specifically, what does it change about which lesson a researcher should actually take from the case?

📄 View solution
Exercise 3

Describe a real or hypothetical situation from your own life or work where you noticed yourself only looking for information that confirmed something you already believed. Explain what a genuine, confirmation-bias-resistant version of that same search would have looked like.

📄 View solution

Chapter 7 Quick Reference

  • Probability sampling gives every population member a known chance of selection; non-probability sampling doesn't, limiting real generalization
  • The 1936 Literary Digest poll (10M mailed, 2.3M returned) wrongly predicted Landon 57%-43%; Roosevelt actually won 62%-38%
  • George Gallup's smaller, properly constructed sample correctly predicted the real result — the founding moment of modern scientific polling
  • Squire's real 1988 reanalysis found non-response bias, not just sample-frame skew, was the larger real driver of the error
  • Peter Wason's 1960 study coined "confirmation bias" — the tendency to seek confirming rather than disconfirming evidence