Statistical Illusions nobody has to lie
ILLUSION 03 · A FORGOTTEN BASE RATE

The Positive

You build a test that sounds excellent. It still cries wolf.

A screening test is judged on two numbers. How often it catches the disease when it is there, called sensitivity, and how often it clears you when it is not, called specificity. A test that scores high on both is, by any ordinary use of the word, a good test.

You are going to build one, point it at a population, and then answer a single question about it. Not whether it is any good. Whether a positive result from it means anything. Those are not the same question, and the gap between them is the whole illusion.

Setup

The brief

Build a screening test you would be happy to call excellent, and decide who to screen with it. You have three dials. To keep them honest, here is what each one means, counted out in people rather than percentages.

Sensitivity

Of every 10 people who have the disease, how many the test correctly flags.

Specificity

Of every 100 people who are healthy, how many the test correctly clears.

Prevalence

Of every 1,000 people you screen, how many actually have the disease.

What is on the table, and what is not

Every number is one you set, in the open. Nothing is hidden now and nothing is revealed later that was withheld from you at the start.

The only thing held back is the crowd your test produces. You will build the test, say what a positive from it is worth, and only then count the people it flags. You will still be surprised.

Operate

Design the test

Set all three dials. Push sensitivity and specificity as high as you can; a real screening test rarely clears the high nineties on both, but you are allowed to. Then choose how common the disease is in the people you point it at.

catches the sick
clears the healthy
how common it is

One thousand people you screened – each dot is one person

set the dials to build your test
has the condition
flagged, but healthy
cleared
Fig. 1 · your screened population, one dot per person. The coral dots up in the corner are the ones who actually have the condition. The flags your test raises are held back until you commit a guess.
Commit

Say what a positive is worth

· unlocks once you design the test

Here is the one question. Take everyone your test flags as positive, the whole pile of them, and ask how many actually have the disease. Not how good the test is. What a single positive result from it is worth. Commit to a number. There is no unpick button, here or in a doctor's office.

50 in 100
One way. The count comes next, and it does not move to spare you.
Your guess

–

Reveal

Now count the ones it flagged

· unlocks when you commit

You built a test almost anyone would call excellent. Most of the people it flags are healthy.

People your test flagged
–
out of the 1,000 you screened
Of those, actually sick
–
the ones the test was for
So a positive is worth
–
the answer to the one question

Gerd Gigerenzer put the textbook version to doctors as a probability: a 1% prevalence, a test that catches 90% of cancers and falsely flags 9% of healthy women. Reframed as counts, it reads:

Ten out of every 1,000 women have breast cancer. Of these 10 women with breast cancer, 9 test positive. Of the 990 women without cancer, about 89 nevertheless test positive.

Nine of the 98 women who test positive have cancer. About 1 in 10.

Repair

So what do you actually do

· unlocks after the reveal

There is a fix, and you just used it without being told. The problem is nearly impossible as three percentages and nearly easy as a count of people, because counting keeps the healthy crowd in view instead of normalising it away. When a result comes back positive, ask for the numbers the way this figure gives them: out of how many, and how many of those are real.

Partly. Be honest about how much.
Debrief

What just happened

· unlocks after the repair

You designed a good test, guessed high about what it was worth, and were wrong by a lot, in the same direction almost everyone is wrong, doctors included. Nothing was rigged. The test did exactly what you set it to do. The population it screened simply had so few sick people in it that the test's rare mistakes outnumbered its real catches.

What was absent from the table

Every number about the test was in front of you. What did not feel like a number was the crowd it was pointed at: the 990 healthy people whose small error rate does all the damage. The base rate is not withheld so much as easy to forget, because it is the part of the problem that is about everyone except the person in front of you. It is the population nobody counted.

Who caught this

Casscells, Schoenberger and Graboys put a version of the question to doctors in 1978 (NEJM 299(18), 999) and most missed it. Kahneman and Tversky named the pattern the same decade (1973, Psychological Review 80(4), 237). Gerd Gigerenzer and Ulrich Hoffrage showed the fix in 1995 (Psychological Review 102(4), 684), and Gigerenzer and colleagues took it to working physicians in 2007 (Psychological Science in the Public Interest 8(2), 53).

One honest caveat

The failure is not that people cannot do arithmetic; it is that the way a question is written decides how many get it right. Counts help a great deal and do not cure it, and base-rate neglect is a tendency, not a law: change the wording and people sometimes lean too hard on the base rate instead. And to be plain, because the subject earns it: nothing here says do not get screened. It says know what a positive result is telling you.