Statistical Illusions nobody has to lie
ILLUSION 01 · MANUFACTURED SIGNIFICANCE

The Forking Paths

You can pull a discovery out of nothing at all.

Below is a study with nothing in it. Not a weak effect, not a small one. Nothing. Forty subjects were drawn from a random number generator in your browser a moment ago, and the difference between the two groups was set to exactly zero.

You are going to find a publishable result in it anyway. Every choice you make on the way there will be one a careful person has defended in print, and I will show you who defended it. The question this answers is not whether it can be done. It is how long it takes you.

Setup

The study

Forty adults. Half took a twenty-minute nap; half sat quietly with a magazine. Afterwards everyone did the same short computer task, and three things were recorded: how fast they responded, how accurately, and how awake they said they felt. Age, sex, and hours slept the night before were noted at intake, as they would be in any real protocol.

Full disclosure, before you start

There is no nap effect in this data. The two groups are drawn from the same distribution and the true difference is exactly zero, by construction.

Nothing is hidden from you here and nothing will be revealed later that was withheld now. The seed is in the URL, so this exact study is reproducible and shareable. Everything below is computed in your browser as you read.

# Group Reaction time Accuracy Alertness Sex Age Slept
showing 8 of 40 subjects

Reaction time in milliseconds (lower is faster) · accuracy in percent · alertness self-rated 1–7. Generated from a seeded normal draw; the three measures correlate at about r = .5, as measures of one construct do.

Operate

Your analysis

Four decisions. None of them is cheating, and none of them is even unusual. Each one has been made, and justified, in the published literature. Turn them until the number at the bottom right of the panel drops below 0.05.

Why this choice is defensible

Hover or select any option above. Every one carries a real, published justification. That is what makes this an accident rather than a fraud.

Your path through the garden – one analysis at a time

p =  0 paths tried
your current analysis
a path you already tried
Nothing here is irreversible. Yet.
Fig. 1 · Welch two-sample t-tests computed live on the forty subjects above. The horizontal axis is the p-value, compressed so the 0.05 convention sits where it does in practice: near the end, and easy to reach.
Commit

Publish it

· unlocks when an analysis crosses p < 0.05

This is the step that separates an exploration from a finding. Once you publish, the other paths stop existing. Not because anyone destroyed them, but because nobody ever hears about them.

One way. There is no unpublish button, here or anywhere.
Published

Reveal

The paths you didn’t take

· unlocks when you publish

You took one path. There were 251.

The figure above has not changed its axis, its scale, or its data. Every other analysis you could have run on these same forty subjects has simply been drawn in alongside yours. Your published finding is the indigo ring.

Paths in your garden
3 measures × 7 subgroups × 3 outlier rules × 4 stopping points
Of those, “significant”
on data with no effect in it
Studies like yours containing at least one significant path
not yet run
One measure only, reproducing the 2011 paper
published figure: 60.7%
Fresh noise each time, same four knobs, computed here. Takes about a second.
Repair

The same move, made honestly

· unlocks after the reveal

Nothing about the four knobs was wrong. What was wrong was choosing among them after seeing the data. So make one change and nothing else: write down which measure, which subgroup, which outlier rule and which stopping point you will use before the data exists, then run the same six hundred studies and look only at that path.

Same generator. Same seeds. Same test. One path each.
Free to choose after the fact
of 600 studies produced a “significant” finding
One path, fixed in advance
nominal false-positive rate: 5%
Debrief

What just happened

· unlocks after the repair

You did deliberately, in about ninety seconds, a thing that usually happens slowly and without anyone noticing. That distinction matters more than it looks. The researchers who described this were careful to separate fishing (running many analyses and reporting the best) from the far more common case: running one analysis, chosen honestly, in a world where different data would have led the same honest person to a different choice. The second one leaves no trace in the paper, in the lab notebook, or in the memory of the person who did it.

What was absent from the table

The paths you didn’t take. Nothing was hidden and nothing was falsified. The other 250 analyses simply never had to be mentioned, because they were never run.

Who caught this

Joseph Simmons, Leif Nelson & Uri Simonsohn: “False-Positive Psychology,” Psychological Science 22(11), 2011, and their continuing work at Data Colada, their research-credibility blog. Andrew Gelman & Eric Loken named the garden of forking paths (2013; American Scientist 102(6), 460, 2014). John Ioannidis modelled the consequences for the literature (PLoS Medicine 2(8), 2005).

One honest caveat

The 2011 paper’s most-quoted rule (require at least twenty observations per cell) was withdrawn by its own authors in 2018, because it “led people to focus on the wrong aspect of disclosure.” Their current answer is preregistration, which is what the repair above actually does. And Ioannidis’s title is a modelling result under stated assumptions, not an audit of the literature; there is a standing peer-reviewed critique.