Statistical Illusions nobody has to lie

Sources

Every primary source behind this anthology, the code behind its numbers, and the search terms the kickers keep out on purpose.

ILLUSION 01 · MANUFACTURED SIGNIFICANCE

The Forking Paths

The exhibit's own Debrief names who caught this and states one honest caveat. Here is the full record behind both.

  1. Simmons, J. P., Nelson, L. D., & Simonsohn, U. (2011). False-Positive Psychology: Undisclosed Flexibility in Data Collection and Analysis Allows Presenting Anything as Significant. Psychological Science, 22(11), 1359–1366.

    doi.org/10.1177/0956797611417632
  2. Gelman, A., & Loken, E. (2013). The Garden of Forking Paths: Why Multiple Comparisons Can Be a Problem, Even When There Is No "Fishing Expedition" or "p-Hacking" and the Research Hypothesis Was Posited Ahead of Time. Unpublished manuscript, Department of Statistics, Columbia University. No journal, volume, or DOI: it has never been formally published.

    sites.stat.columbia.edu/gelman/research/unpublished/p_hacking.pdf
  3. Gelman, A., & Loken, E. (2014). The Statistical Crisis in Science. American Scientist, 102(6), 460. The peer-reviewed version of the same argument, in plainer language.

    americanscientist.org/article/the-statistical-crisis-in-science
  4. Ioannidis, J. P. A. (2005). Why Most Published Research Findings Are False. PLoS Medicine, 2(8), e124. A modelling result under stated assumptions, not an audit of the literature. Goodman & Greenland's peer-reviewed critique (PLoS Medicine, 2007, 4(4), e168) argues the model cannot prove most published claims are false; Ioannidis replied the same year.

    doi.org/10.1371/journal.pmed.0020124
  5. Data Colada. The research-credibility blog run by the same three authors, publishing since 2013.

    datacolada.org
What the literature calls this

The kicker above says manufactured significance, not p-hacking. The term is common, but its own originators are wary of it. Simmons, Nelson and Simonsohn call it researcher degrees of freedom and never write p-hacking themselves. Gelman and Loken, who named the garden of forking paths, later wrote that they regret how fast fishing and p-hacking spread, because both words imply someone did this on purpose. Multiple comparisons is the older, broader statistical term for the same family of problem. The concept belongs to the researchers above; none of them coined the word.

The interaction it borrows

The mechanic, choosing an analysis by hand and watching the p-value move, adapts FiveThirtyEight's Hack Your Way to Scientific Glory (Christie Aschwanden & Ritchie King, 2015). Their working example was partisan; this exhibit keeps the interaction and swaps in Simmons, Nelson and Simonsohn's own study instead.

ILLUSION 02 · A HIDDEN VARIABLE

The Reversal

The exhibit's Debrief names who caught this and states one honest caveat. Here is the full record behind both.

  1. Simpson, E. H. (1951). The Interpretation of Interaction in Contingency Tables. Journal of the Royal Statistical Society: Series B, 13(2), 238–241. The paper that sets out the reversal, and relabels its own table to show the correct answer depends on the story, not the arithmetic.

    doi.org/10.1111/j.2517-6161.1951.tb00088.x
  2. Blyth, C. R. (1972). On Simpson's Paradox and the Sure-Thing Principle. Journal of the American Statistical Association, 67(338), 364–366. The paper that gave the effect its name.

    doi.org/10.1080/01621459.1972.10482387
  3. Pearl, J. (2014). Comment: Understanding Simpson's Paradox. The American Statistician, 68(1), 8–13. Which number is correct is a causal question, not a statistical one, and sometimes the answer is in neither table.

    doi.org/10.1080/00031305.2014.876829
  4. Julious, S. A., & Mullee, M. A. (1994). Confounding and Simpson's paradox. BMJ, 309(6967), 1480–1481. The diabetic-cohort version, where the reversal changes a clinical reading and survives without binning the confounder.

    doi.org/10.1136/bmj.309.6967.1480
  5. Baseball-Reference. Standard batting, David Justice (justida01) and Derek Jeter (jeterde01), 1995–1997. The twelve numbers behind the figure are read off these pages.

    baseball-reference.com/players/j/jeterde01.shtml
  6. Ross, K. A. (2004). A Mathematician at the Ballpark. Pi Press. The popularizer of the Justice-Jeter example; the data itself is Baseball-Reference's.

Attribution, precisely

Simpson never uses the word "paradox" and never uses the word "causal"; Blyth named it in 1972, a reading that rests on Pearl's history of it. Earlier reports of the vanishing association go back to Pearson (1899) and Yule (1903), and sign reversal to Cohen and Nagel (1934). Pearl's account of when to trust which table is argued, not universally settled.

The example, and why baseball

The exhibit uses baseball because it needs no causal claim: nobody argues that at-bats cause hits in a way that must be adjusted for, so the reversal is transparently an artifact of weighting. The medical versions each drag a causal argument along, and the famous kidney-stone table (Charig et al., 1986) also carries two traps: "group 2" is a composite severity category rather than a size, and the two treatments ran in non-overlapping eras. They are kept off the page on purpose.

ILLUSION 03 · A FORGOTTEN BASE RATE

The Positive

The exhibit's Debrief names who caught this and states one honest caveat. Here is the full record behind both, including the honest effect size and the real-world numbers.

  1. Casscells, W., Schoenberger, A., & Graboys, T. B. (1978). Interpretation by Physicians of Clinical Laboratory Results. New England Journal of Medicine, 299(18), 999–1001. The original doctors-and-Bayes result. Framed carefully: the problem never stated the sensitivity, so the finding is that the format of a question drives the answer, not that physicians cannot reason. The spelling is Graboys, though the literature often prints "Grayboys."

    doi.org/10.1056/NEJM197811022991808
  2. Kahneman, D., & Tversky, A. (1973). On the Psychology of Prediction. Psychological Review, 80(4), 237–251. Where base-rate neglect is named, alongside the finding that it is a tendency, not a law: change the wording and people sometimes lean too hard on the base rate instead.

    doi.org/10.1037/h0034747
  3. Gigerenzer, G., & Hoffrage, U. (1995). How to Improve Bayesian Reasoning Without Instruction: Frequency Formats. Psychological Review, 102(4), 684–704. The repair: restating probabilities as natural frequencies, counts sharing one reference class.

    doi.org/10.1037/0033-295X.102.4.684
  4. Gigerenzer, G., Gaissmaier, W., Kurz-Milcke, E., Schwartz, L. M., & Woloshin, S. (2007). Helping Doctors and Patients Make Sense of Health Statistics. Psychological Science in the Public Interest, 8(2), 53–96. The mammography problem and the 160 gynaecologists. Volume 8(2) is 2007, though its DOI string reads 2008; both are correct.

    doi.org/10.1111/j.1539-6053.2008.00033.x
  5. McDowell, M., & Jacobs, P. (2017). Meta-Analysis of the Effect of Natural Frequencies on Bayesian Reasoning. Psychological Bulletin, 143(12), 1273–1312. The honest effect size, 226 estimates across 35 articles: 24% correct with counts against 4% with probabilities.

    doi.org/10.1037/bul0000126
  6. Lehman, C. D., et al. (2017). National Performance Benchmarks for Modern Screening Digital Mammography. Radiology, 283(1), 49–58. The real-world number behind the debrief: across 1.6 million screening mammograms, a positive meant cancer about 4% of the time.

    doi.org/10.1148/radiol.2016161174
The one guardrail

This exhibit is about what a positive result means, never about whether to screen. Screening intervals are a live clinical-policy question the piece has no business in. Nothing here says do not get screened; it says know what a positive result is telling you.

Numbers used where, and why two sets

The interactive uses your own dials. The quoted study uses Gigerenzer's published wording, a 1% prevalence, 90% sensitivity and a 9% false-positive rate, which lands on his 9 of 98. The real-world 4% is the Breast Cancer Surveillance Consortium's measured figure. Where the textbook and the clinic disagree, both are named rather than blended into one tidy number.

ILLUSION 04 · AN UNCOUNTED POPULATION

The Missing Planes

The exhibit's Debrief names who caught this and states one honest caveat; the essay's closing panel separates the famous story from the record. Here is the full record behind both.

  1. Wald, A. (1980). A Reprint of "A Method of Estimating Plane Vulnerability Based on Damage of Survivors." CRC 432. Center for Naval Analyses, reprinting the 1943 memoranda. Pure mathematics and numeric tables; no damage diagram anywhere in it. The short title is the one on the reprint; the circulating "of Enemy Fire" is not.

    cna.org/analyses/1980/a-method-of-estimating-plane-vulnerability
  2. Wallis, W. A. (1980). The Statistical Research Group, 1942–1945. Journal of the American Statistical Association, 75(370), 320–330. The insider account of the group Wald worked in; its two vague mentions of the armor problem do not name him.

    doi.org/10.1080/01621459.1980.10477469
  3. Mangel, M., & Samaniego, F. J. (1984). Abraham Wald's Work on Aircraft Survivability, with comment by J. O. Berger and rejoinder. Journal of the American Statistical Association, 79(386), 259–271. The worked example in the exhibit reaches us through this rejoinder, which carries Berger's caveat and the "we do not know whether it was used during World War II" line.

    doi.org/10.1080/01621459.1984.10478038
  4. Casselman, B. (2016). The Legend of Abraham Wald. AMS Feature Column, June 2016. American Mathematical Society. The source for what does and does not survive in the record, including the "no diagram" finding above.

    ams.org/publicoutreach/feature-column/fc-2016-06
Why the map is drawn, not found

No damage diagram was ever published, so the bullet map in the exhibit is a browser simulation from a visible seed, and it says so. The often-reproduced bullet-riddled bomber is a 2016 Wikipedia illustration of a Lockheed PV-1 Ventura, an aircraft in none of Wald's examples; the exhibit draws its own generic plane, which avoids both the mislabelling and a share-alike licence.

Getting both directions right

The story overstates Wald and the usual correction understates him. He mentions armor once, in general terms, calling his estimates something that "can be used as guides for locating protective armor," with the engine result coming from a hypothetical worked example. The honest caveat is Berger's, which Wald accepted: the method assumes a constant per-hit survival probability that survivor-only data cannot test. Documented use came later, in Vietnam and on the B-52, not the Second World War.

The code behind these numbers

No dataset and no server sit behind any figure on this site, including the ones above. Every number is a few dozen constants and about a hundred and thirty lines of statistics, run in your browser while you read, and all of it is public.

github.com/dustincole-data/statistical-illusions →