The Positive
You build a test that sounds excellent. It still cries wolf.
A screening test is judged on two numbers. How often it catches the disease when it is there, called sensitivity, and how often it clears you when it is not, called specificity. A test that scores high on both is, by any ordinary use of the word, a good test.
You are going to build one, point it at a population, and then answer a single question about it. Not whether it is any good. Whether a positive result from it means anything. Those are not the same question, and the gap between them is the whole illusion.
The brief
Build a screening test you would be happy to call excellent, and decide who to screen with it. You have three dials. To keep them honest, here is what each one means, counted out in people rather than percentages.
Of every 10 people who have the disease, how many the test correctly flags.
Of every 100 people who are healthy, how many the test correctly clears.
Of every 1,000 people you screen, how many actually have the disease.
Every number is one you set, in the open. Nothing is hidden now and nothing is revealed later that was withheld from you at the start.
The only thing held back is the crowd your test produces. You will build the test, say what a positive from it is worth, and only then count the people it flags. You will still be surprised.
Design the test
Set all three dials. Push sensitivity and specificity as high as you can; a real screening test rarely clears the high nineties on both, but you are allowed to. Then choose how common the disease is in the people you point it at.
One thousand people you screened – each dot is one person
Say what a positive is worth
· unlocks once you design the testHere is the one question. Take everyone your test flags as positive, the whole pile of them, and ask how many actually have the disease. Not how good the test is. What a single positive result from it is worth. Commit to a number. There is no unpick button, here or in a doctor's office.
–
Now count the ones it flagged
· unlocks when you commitYou built a test almost anyone would call excellent. Most of the people it flags are healthy.
Gerd Gigerenzer put the textbook version to doctors as a probability: a 1% prevalence, a test that catches 90% of cancers and falsely flags 9% of healthy women. Reframed as counts, it reads:
Ten out of every 1,000 women have breast cancer. Of these 10 women with breast cancer, 9 test positive. Of the 990 women without cancer, about 89 nevertheless test positive.
Nine of the 98 women who test positive have cancer. About 1 in 10.
So what do you actually do
· unlocks after the revealThere is a fix, and you just used it without being told. The problem is nearly impossible as three percentages and nearly easy as a count of people, because counting keeps the healthy crowd in view instead of normalising it away. When a result comes back positive, ask for the numbers the way this figure gives them: out of how many, and how many of those are real.
Partly. The best measurement of the effect, a meta-analysis of 226 estimates across 35 studies, puts the share of people who get it right at 24% with counts against 4% with probabilities. That is a six-fold improvement and a genuinely large one, and three quarters of people still get it wrong even in the easy format. Thinking in counts is a real help, not a cure.
And the textbook numbers flatter the test. Across 1.6 million real screening mammograms, the measured chance that a positive means cancer was about 4%, closer to 1 in 23 than the 1 in 10 the classroom version gives. The gap runs the safe way here, but it is the same lesson: a positive on a rare condition is usually a false alarm, and the better you know the base rate, the less a single result should move you.
What just happened
· unlocks after the repairYou designed a good test, guessed high about what it was worth, and were wrong by a lot, in the same direction almost everyone is wrong, doctors included. Nothing was rigged. The test did exactly what you set it to do. The population it screened simply had so few sick people in it that the test's rare mistakes outnumbered its real catches.
Every number about the test was in front of you. What did not feel like a number was the crowd it was pointed at: the 990 healthy people whose small error rate does all the damage. The base rate is not withheld so much as easy to forget, because it is the part of the problem that is about everyone except the person in front of you. It is the population nobody counted.
Casscells, Schoenberger and Graboys put a version of the question to doctors in 1978 (NEJM 299(18), 999) and most missed it. Kahneman and Tversky named the pattern the same decade (1973, Psychological Review 80(4), 237). Gerd Gigerenzer and Ulrich Hoffrage showed the fix in 1995 (Psychological Review 102(4), 684), and Gigerenzer and colleagues took it to working physicians in 2007 (Psychological Science in the Public Interest 8(2), 53).
The failure is not that people cannot do arithmetic; it is that the way a question is written decides how many get it right. Counts help a great deal and do not cure it, and base-rate neglect is a tendency, not a law: change the wording and people sometimes lean too hard on the base rate instead. And to be plain, because the subject earns it: nothing here says do not get screened. It says know what a positive result is telling you.