Tags: statistics concept

P-Hacking

Date: 2026-08-16


Adjusting the analysis until the result crosses the line. It rarely feels like cheating from the inside — each individual choice is defensible, and the cumulative effect is that you can find significance in noise essentially whenever you want to.


What it is

P-hacking is exploiting the flexibility in analysis choices to produce a significant result, then reporting it as though the analysis had been fixed in advance.

The defining property: the choices were made after seeing the data. The same analysis, decided beforehand, would be entirely legitimate.

The moves

Each of these is a normal analytical decision. Made after looking, each one is a p-hack:

MoveSounds like
Stop when it’s significant”It reached significance, so we ended it” — Peeking
Keep running when it isn’t”We gave it more time to settle”
Try several metrics, report the winner”Revenue per visitor also moved” — The Multiple Comparisons Problem
Find a segment that worked”It won on mobile” — Segmentation (test results)
Exclude outliers after seeing them”That £5,000 order distorted it”
Exclude a bad week”There was an outage on the Tuesday”
Switch to one-tailed”We only ever cared about improvement”
Change the primary metric”Conversion was always the real goal”

Any one is arguable. Two or three together will find significance in pure noise most of the time.

Why the p-value stops meaning anything

A p-value is the probability of data this extreme given the null and given the procedure you committed to. It’s a property of the procedure, not of the numbers.

committed procedure    one metric, one endpoint    → p means what it says
flexible procedure     stop when significant,      → p means nothing,
                       pick the best metric,         and still prints as 0.03
                       drop the bad week

In plain terms: the printed p-value assumes you did exactly one thing. If you tried six and reported the best, the number on the screen is describing an experiment you didn’t run.

Deliberate versus accidental

P-hacking implies intent, and most of it isn’t. The same statistical damage arrives through entirely good-faith reasoning — noticing that an outlier “obviously” distorted things, or that a segment “makes sense” — with no awareness that a fork was taken.

That’s The Garden of Forking Paths, and it’s the far more common version. Distinguishing them matters for how you respond: the deliberate version needs an ethical answer, the accidental version needs a procedural one.

The procedural answer

Remove the degrees of freedom before the data exists. Everything else is exhortation.

Recognising it in someone else’s result

Questions worth asking, none of them accusatory:

  • What was the primary metric, and where is that written down?
  • How many metrics were examined?
  • Was the endpoint set in advance?
  • Were the segments pre-registered?
  • Were any exclusions applied, and when were they decided?

A result that can’t answer those is not necessarily wrong — it’s unverifiable, which is a different and more useful thing to say than “I don’t believe it”.

Why it matters commercially

The failure isn’t academic embarrassment. It’s that a p-hacked winner gets shipped, doesn’t work, and nobody connects the two. The programme’s measured win rate stays high while its actual contribution doesn’t, and the gap only becomes visible in a holdout — usually a year later, if anyone runs one.