Paradox & Proof Entry #0173 Classified Declassified

Why the sample that looks representative usually is not

A sample that resembles the population on visible traits was selected for resemblance. Selection on the visible says nothing about the rest.

No visual record attached The written record below is complete.
Plate 435 — The margin that was never in the table

Intuition test — answer before you read on

Why does demographic representativeness fail to guarantee a valid sample?

A panel matches the census on age, region and gender to within a point. It is also composed entirely of people who agree to answer surveys for a small fee, and the thing being measured correlates with exactly that willingness.

What everyone sees

Demographic match is read as representativeness, because demographics are what can be checked. Matching on observable margins is genuinely useful and genuinely insufficient. It constrains the variables used for weighting and leaves every unmeasured variable free, including the one the study is about.

What is actually happening

Tversky and Kahneman documented the underlying error as belief in the law of small numbers: people expect samples to resemble the population in all respects, treating similarity on salient features as evidence of similarity overall, and consequently underestimating how much a small or self-selected sample can diverge. Bethlehem’s analysis of web survey selection sets out the structural version: where participation is voluntary, bias depends on the correlation between the propensity to participate and the variable being measured, and demographic weighting corrects only the part of that correlation running through the weighting variables. So a panel can match the population precisely on everything recorded and remain badly wrong about the thing it was built to measure.

Why it stays hidden

The gap hides because the visible match is published and the invisible mismatch cannot be. Methodology sections report the demographic profile, which is checkable, and are silent about participation propensity, which is not. A reader who scrutinises the available evidence will find it reassuring, and their scrutiny cannot reach the variable that matters.

Matching on what can be checked constrains only what was checked. The variable that decides the result is the one no weighting touched.

Matching on what can be checked constrains only what was checked. The variable that decides the result is the one no weighting touched.

The hidden part — entry #0173

Collect this card

Matching on what can be checked constrains only what was checked. The variable that decides the result is the one no weighting touched.

0 / 10,000 collected

Sources & further reading 2
  1. Tversky & Kahneman — belief in the law of small numbers
  2. Bethlehem — selection bias in web surveys

Circulate this file

Annotations are reserved for archive members.

Sign in to annotate