Why a ninety-nine per cent accurate test is usually wrong
A test that is right ninety-nine times in a hundred can still be wrong about most of the cases it flags, and simple arithmetic shows why.
Filed by The Archivist 2 min read
Intuition test — answer before you read on
A test with ninety-nine per cent accuracy looks for something present in one case per thousand. Roughly how often is a flag correct?
Correct answer: B
Roughly one flag in eleven: ten true cases per ten thousand against about a hundred false alarms drawn from the rest. The lesson is arithmetic, not a verdict on any particular instrument — rare targets guarantee that most alerts are false however good the test.
A fraud detector is described as ninety-nine per cent accurate. It flags an account. The intuitive conclusion is that the account is almost certainly fraudulent. Where fraud is rare, the opposite is nearer the truth, and the detector has made no error at all.
What everyone sees
Accuracy is heard as the probability of being right about this case. Ninety-nine per cent accurate is taken to mean ninety-nine per cent likely correct whenever the instrument speaks, which makes the rarity of the target feel like a technicality rather than half of the calculation.
What is actually happening
Work it through with ten thousand accounts, ten of them fraudulent. The detector catches about ten of the ten and wrongly flags about one per cent of the remaining 9,990, which is roughly a hundred. So about a hundred and ten accounts are flagged and about ten of those are genuine cases: fewer than one in ten. Nothing about the instrument changed, only the rarity of what it looks for.
Why it stays hidden
The base rate is missing from the sentence that reaches the decision-maker. Accuracy is one number and travels well; prevalence is a second number held by somebody else, so the two rarely meet. Studies of this arithmetic have found trained professionals reaching the intuitive answer under time pressure, and the error largely disappears when the same problem is stated in counts instead of percentages.
Accuracy is a property of the test. Whether an alert is right also depends on how rare the thing is.
Accuracy is a property of the test. Whether an alert is right also depends on how rare the thing is.
Collect this card
Accuracy is a property of the test. Whether an alert is right also depends on how rare the thing is.
0 / 10,000 collected
Sources & further reading 3
- Casscells, Schoenberger & Graboys, "Interpretation by Physicians of Clinical Laboratory Results", New England Journal of Medicine, 1978
- Eddy, "Probabilistic Reasoning in Clinical Medicine", in Judgment under Uncertainty, 1982
- Gigerenzer & Hoffrage, "How to Improve Bayesian Reasoning Without Instruction: Frequency Formats", Psychological Review, 1995
Cross-references
Related files
Filed near this one in the index.
-
No visual on fileStatistical Illusions Entry #0437
The reason a study that replicates is worth two that do not
A single study is a claim. A replication is a test of that claim. Two unreplicated studies are two untested claims, not twice the evidence.
AdeptThe hidden part #0437Novelty fills journals. Replication fills knowledge. The incentive structure rewards the wrong one.
Statistics Open file -
No visual on fileStatistical Illusions Entry #0436
Why a p-value is not a probability that you are right
The p-value measures how surprising the data would be if the null hypothesis were true. It says nothing about the probability that the hypothesis itself is true.
NoviceThe hidden part #0436A p-value tells you how surprising the data is. It does not tell you how true the hypothesis is.
Statistics Open file -
No visual on fileScarcity & Queues Entry #0461
Why a paywall creates a different reader
The paywall does not merely filter who reads. It changes how the remaining readers read, because the cost they paid becomes part of the experience.
NoviceThe hidden part #0461The paywall does not select better readers. It makes the same readers better at reading.
Scarcity Open file -
No visual on fileScarcity & Queues Entry #0457
Why a members-only door raises the perceived quality inside
A restriction implies that something is being protected. The inference is made before anything inside has been seen, and it survives finding nothing there.
NoviceThe hidden part #0457Restriction and value are separable. The door can be installed first, and the perceived quality assembles itself behind it.
Scarcity Open file
Annotations are reserved for archive members.
Sign in to annotate