Statistical Illusions Entry #0437 Classified Declassified

The reason a study that replicates is worth two that do not

A single study is a claim. A replication is a test of that claim. Two unreplicated studies are two untested claims, not twice the evidence.

No visual record attached The written record below is complete.
Plate 240 — The test the incentive forgot

Intuition test — answer before you read on

Why does one replicated study carry more weight than two separate unreplicated studies?

Two separate laboratories run the same experiment. One finds a significant effect; the other finds a null result. A third laboratory replicates the first laboratory’s protocol and finds the same significant effect. The replicated finding now carries more weight than either single study did alone, and the non-replication is informative rather than damaging.

What everyone sees

The temptation is to count studies: two positives versus one negative means the effect is probably real. But counting treats all studies as interchangeable units of evidence, ignoring that a replication tests the specific claim of the first study while a second original study may differ in ways that make comparison unreliable.

What is actually happening

Open Science Collaboration’s large-scale replication project found that only about 36 per cent of psychology studies replicated at the original effect size. Ioannidis argued that most published research findings are false, partly because single underpowered studies with publication bias inflate the rate of false positives. A successful replication — same protocol, independent lab, preregistered analysis — provides a fundamentally different kind of evidence from a second original study, because it controls for the degrees of freedom that allowed the first finding to emerge by chance or by analytic flexibility.

Why it stays hidden

The asymmetry hides because the publishing system treats all significant results as equivalent contributions. A replication of an existing finding is often considered less publishable than a novel finding, which means the most valuable kind of evidence is the least rewarded. The result is a literature full of unreplicated claims, each treated as a brick in the wall of knowledge, when they are actually provisional hypotheses awaiting a test that the incentive structure discourages.

Novelty fills journals. Replication fills knowledge. The incentive structure rewards the wrong one.

Novelty fills journals. Replication fills knowledge. The incentive structure rewards the wrong one.

The hidden part — entry #0437

Collect this card

Novelty fills journals. Replication fills knowledge. The incentive structure rewards the wrong one.

0 / 10,000 collected

Sources & further reading 2
  1. Open Science Collaboration — estimating the reproducibility of psychological science
  2. Ioannidis — why most published research findings are false

Circulate this file

Annotations are reserved for archive members.

Sign in to annotate