Tags: statistics concept
P-Hacking
Date: 2026-08-16
Adjusting the analysis until the result crosses the line. It rarely feels like cheating from the inside — each individual choice is defensible, and the cumulative effect is that you can find significance in noise essentially whenever you want to.
What it is
P-hacking is exploiting the flexibility in analysis choices to produce a significant result, then reporting it as though the analysis had been fixed in advance.
The defining property: the choices were made after seeing the data. The same analysis, decided beforehand, would be entirely legitimate.
The moves
Each of these is a normal analytical decision. Made after looking, each one is a p-hack:
| Move | Sounds like |
|---|---|
| Stop when it’s significant | ”It reached significance, so we ended it” — Peeking |
| Keep running when it isn’t | ”We gave it more time to settle” |
| Try several metrics, report the winner | ”Revenue per visitor also moved” — The Multiple Comparisons Problem |
| Find a segment that worked | ”It won on mobile” — Segmentation (test results) |
| Exclude outliers after seeing them | ”That £5,000 order distorted it” |
| Exclude a bad week | ”There was an outage on the Tuesday” |
| Switch to one-tailed | ”We only ever cared about improvement” |
| Change the primary metric | ”Conversion was always the real goal” |
Any one is arguable. Two or three together will find significance in pure noise most of the time.
Why the p-value stops meaning anything
A p-value is the probability of data this extreme given the null and given the procedure you committed to. It’s a property of the procedure, not of the numbers.
committed procedure one metric, one endpoint → p means what it says
flexible procedure stop when significant, → p means nothing,
pick the best metric, and still prints as 0.03
drop the bad week
In plain terms: the printed p-value assumes you did exactly one thing. If you tried six and reported the best, the number on the screen is describing an experiment you didn’t run.
Deliberate versus accidental
P-hacking implies intent, and most of it isn’t. The same statistical damage arrives through entirely good-faith reasoning — noticing that an outlier “obviously” distorted things, or that a segment “makes sense” — with no awareness that a fork was taken.
That’s The Garden of Forking Paths, and it’s the far more common version. Distinguishing them matters for how you respond: the deliberate version needs an ethical answer, the accidental version needs a procedural one.
The procedural answer
Remove the degrees of freedom before the data exists. Everything else is exhortation.
- Pre-Registration. One primary metric, one endpoint, named segments, named exclusion rules, a decision rule. Committed somewhere edits are visible
- Fixed sample and endpoint, or a design built for flexible stopping — Sequential Testing, Always-Valid Inference
- Exclusion rules from historical data, so the cap is set before you know who it excludes — Winsorisation and Capping
- Post-hoc findings labelled as hypotheses, in their own section, carrying no decision — Experiment Archive
Recognising it in someone else’s result
Questions worth asking, none of them accusatory:
- What was the primary metric, and where is that written down?
- How many metrics were examined?
- Was the endpoint set in advance?
- Were the segments pre-registered?
- Were any exclusions applied, and when were they decided?
A result that can’t answer those is not necessarily wrong — it’s unverifiable, which is a different and more useful thing to say than “I don’t believe it”.
Why it matters commercially
The failure isn’t academic embarrassment. It’s that a p-hacked winner gets shipped, doesn’t work, and nobody connects the two. The programme’s measured win rate stays high while its actual contribution doesn’t, and the gap only becomes visible in a holdout — usually a year later, if anyone runs one.