Statistical Illusions
Nobody has to lie for a number to mislead you. Here, you do it yourself.
Darrell Huff called his 1954 book How to Lie with Statistics, and seventy years on it is still the one everybody has heard of. The title is the trouble with it. Lying is something a person decides to do, so if that were the problem you could go looking for the person and ask what they stood to gain. Huff called it lying. Usually nobody is, and that is what makes it hard to catch.
Most writing on this works the same way. Here is a chart that fooled you, here is the trick, now you know better. It is a comfortable thing to read and I do not think it teaches much, because you cannot really feel a mistake that somebody who already knows the answer is describing to you. So this is built the other way around. Every illusion here is one you manufacture yourself, out of numbers you are shown in full before you touch them. Nothing is held back and nothing gets revealed later that was hidden from you at the start. You will still get there.
That works because of the second half of the thesis, which is the half the title misses. Nobody has to lie for a number to mislead you. The defaults do it, and you would have picked the same ones. In the first illusion every option in front of you is a choice somebody has already made and defended in print, and you get the citation at the moment you make it. The claim is not that researchers are crooks. The claim is worse than that, and harder to fix: the careful, defensible, well justified choice is enough on its own.
Four illusions follow. One is playable today; the rest are being built.
1. The Forking Paths
You run a small study. Forty adults, half of them take a twenty minute nap and half sit quietly with a magazine, and afterwards everyone does the same short computer task. There is nothing in it. Both groups came out of the same random number generator and the true difference between them is exactly zero, which you are told before you touch anything.
Then you analyze it. Which of the three things you measured is the outcome? Do you keep the subject who nodded off at the keyboard, or is that an outlier? Everyone, or only the sleep deprived people, where a nap has room to do something? Did you stop at forty, or at thirty, when the pattern already looked clear? None of those questions has a wrong answer, and every answer to them has a defence in the published literature.
Below is one such study with every one of those analyses run at once. Each dot is one analysis, placed by its p-value, and the dashed line is the 0.05 convention.
The garden of forking paths – every analysis of one study with nothing in it
Nobody runs all of them, and that is the part worth sitting with. One analysis is enough, picked honestly, after the data came in, by a person who would have picked differently if the data had come in differently. None of that shows up in the paper. There is no line in the method section for the analyses you never had a reason to run.
2. The Reversal
Two hitters. One of them had the better batting average in 1995. He had it again in 1996, and again in 1997. Then you add the three seasons together and he loses. Nothing was falsified and no season was cherry-picked; the totals are weighted by how many times each man came to the plate, and that number was never put in front of you. You will re-weight it yourself, name a winner, and watch the winner change. What was absent from the table: the variable nobody split by.
3. The Positive
You are going to design a screening programme. You choose how common the disease is, how often the test catches it, and how often it cries wolf, and every number you pick will look reasonable. Then one person tests positive and you commit to a guess: what are the odds they are sick? Most people guess high. So do most doctors. Then you count the thousand people your programme actually touched, one at a time, and go looking for that person in a crowd of false alarms. What was absent from the table: the population nobody counted.
4. The Missing Planes
Bombers come back from Europe with holes in them, and you have armor enough for one part of the plane. The wings are riddled. The fuselage is riddled. The engines are almost clean. You put the armor where the damage is, because that is where the damage is, and then you fly the fleet and watch what happens to it. The answer has been told for eighty years as a story about a clever statistician, and most of the details in that telling are not in the record, which is the same lesson twice. What was absent from the table: the planes that never came home.
What was absent
Four illusions, four subjects with nothing in common, one mechanism. Every one of them is caused by something that was not on the table. The paths you did not take. The variable you did not split by. The population you did not count. The planes that did not come home.
You can audit a number that is in front of you. Check its arithmetic, chase its source, ask who paid for it, ask what it would have been under a different definition. There is no equivalent move for a number that is not there, because there is nothing to point at and nobody to ask. Absence does not look like anything. That is the whole reason nobody has to lie.
Manufacturing that first result takes about ninety seconds, and every choice along the way has a published defence sitting next to it. The number at the end is not the point. The point is how little friction there was between an honest question and a wrong answer.
Handing the question to a language model costs about the same. You type it, a paragraph comes back, and the paragraph is fluent and specific and has no visible seams in it. Whatever was considered and not considered on the way to it is absent from the table in the same way the other two hundred and fifty analyses were, except that this time there is no way to draw the rest of the garden.
I want to be plain about what that is. It is an analogy. Nothing on this site measures a language model, tests one, or shows anything at all about how often they are wrong. All it can do is hand you the feeling once, in a case small enough that the whole garden fits on one axis. Where else you have felt it is your call, not mine.
The planes that came back
You may already know the fourth illusion in its famous form. Bombers return from raids with their wings and fuselage full of holes. Someone proposes armoring the parts with the most holes. A statistician named Abraham Wald says no, armor the engines, because the planes that were hit in the engines are not in the room to be counted.
The mechanism in that story is real, and it is Wald’s. Most of the rest of the telling is not.
The storyThe diagram of a bomber with the bullet holes clustered on the wings.
The recordDrawn in 2016 by a Wikipedia editor and redrawn in 2021. The aircraft in it is a Lockheed PV-1 Ventura, which appears in none of Wald’s examples. His memoranda carry no such diagram; they are pure mathematics and numeric tables.
The story“Not so fast,” Wald told the officers.
The recordThere is no record of the meeting, the officers, or the line. Casselman, who went looking: “There is extremely little source material for what Wald had to say about aircraft damage.” What survives is the memoranda themselves and two vague mentions in W. Allen Wallis’s memoir of the Statistical Research Group, neither of which names Wald.
The storySo he told them to armor the engines.
The recordWald mentions armor once, in general terms, as something his estimates “can be used as guides for locating.” The engine result comes out of a worked example he labels in the same sentence as hypothetical data. He makes no recommendation to any commander. The opposite correction, that he never mentioned armor at all, is wrong too.
The storyAnd it saved the bombers.
The recordMangel and Samaniego, in a peer-reviewed rejoinder: “We do not know whether it was used during World War II.” The documented uses come later, in Vietnam and on the B-52.
Wald, A Method of Estimating Plane Vulnerability Based on Damage of Survivors (CRC 432, 1980, reprinting the 1943 memoranda) · Mangel & Samaniego, JASA 79(386), 1984 · Casselman, The Legend of Abraham Wald, AMS Feature Column, 2016.
None of this makes Wald smaller. In 1943, working only from the aircraft that came back, he built a way to estimate damage he could not observe, and he was careful to say that the method rests on an assumption survivor data cannot check. The caution is the part that did not survive. What got repeated is the version with a good line in it.
The story about the planes that didn’t come back is itself a story that came back.