Tags: statistics concept

The Garden of Forking Paths

Date: 2026-08-16


The same statistical damage as p-hacking, with none of the intent. You made one analysis and never realised how many others you’d have made if the data had come out differently — and it’s those unmade choices that invalidate the p-value.


What it is

The garden of forking paths describes analysis where reasonable, data-dependent choices were made honestly — but where different data would have led to different choices, each equally reasonable.

The key insight: you don’t have to run multiple analyses to suffer multiplicity. It’s enough that you would have.

the data came out this way
        │
        ▼
"mobile looks interesting" ──▶ analyse mobile ──▶ significant ──▶ report

but had it come out otherwise:
   "desktop looks interesting"     → analyse desktop
   "new visitors look interesting" → analyse new visitors
   "the second week looks odd"     → analyse by week
   "that outlier is distorting it" → exclude and re-run

five paths existed. You walked one. The p-value
assumes there was only ever one.

Why this is worse than p-hacking

Not morally — practically.

P-hacking is detectable and correctable. Someone deliberately trying twenty metrics knows they tried twenty, and can be asked.

Forking paths is invisible from the inside. The analyst made one decision, in good faith, that they can fully justify. There’s nothing to confess and nothing to correct, because from where they’re standing there was only ever one analysis.

In plain terms: you can’t count the comparisons you’d have made in a world you didn’t end up in — but the maths counts them anyway.

Where it happens

  • “That segment looks interesting” — the segment was chosen because it looked interesting, which means it was chosen by the data
  • “That week was clearly anomalous” — noticed because it was extreme, which is what noise looks like
  • “This outlier is obviously distorting the result” — obvious only because you saw it
  • “The metric we really care about is…” — clarified after seeing which one moved
  • “It makes sense that mobile would respond differently” — an explanation generated to fit a finding, which would have been generated just as fluently for desktop
  • Choosing a date range that “looks representative”

The signature is a defensible reason arrived at after the observation, for a choice that would have been made differently under different data.

The tell

One question, and it’s the most useful diagnostic here:

Would I have made this same choice if the data had come out the other way?

If yes, the choice was independent of the data and is legitimate. If no, you’ve walked a fork, and the p-value doesn’t account for the branches you didn’t take.

A second version for reading someone else’s work: would they have told me about this segment if it hadn’t been significant? If not, it’s a fork.

The response

The same procedural fix as P-Hacking, and for the same reason — you cannot solve this with care, because care is what produced it:

  • Pre-Registration. Choices fixed before data exists. The only thing that genuinely removes forks
  • A stated analysis plan, including exclusions and segments, however brief
  • Treat post-hoc findings as hypotheses, explicitly labelled and carrying no decision — Segmentation (test results)
  • Replicate anything surprising. A finding that survives a dedicated follow-up test is real; nothing else settles it
  • Correct where multiplicity is unavoidable — Bonferroni and False Discovery Rate

The honest position

This is not a criticism of anyone’s rigour. It’s a structural property of flexible analysis, and it applies to careful, honest, experienced people exactly as much as to careless ones — arguably more, because experience supplies more plausible post-hoc explanations.

Which is why the answer is never “be more careful”. It’s to fix the decisions before the data can influence them.