Tags: statistics concept
Null and Alternative Hypotheses
Date: 2026-08-16
The null is the boring explanation you’re trying to rule out. It’s always a specific claim — usually “the difference is exactly zero” — because you can only compute probabilities against something specific.
The pair
- Null hypothesis (H₀) — no effect. The two groups are samples from the same population, and any difference you see is sampling noise
- Alternative hypothesis (H₁) — there is an effect. What you’d conclude if the null is rejected
For the running example:
H₀ : p_variant = p_control the new checkout changes nothing
H₁ : p_variant ≠ p_control the new checkout changes something
The null has to be precise, the alternative doesn’t. That asymmetry isn’t stylistic — you need one exact number to build a distribution around, and “no difference” gives you one. “Some difference” gives you infinitely many, so there’s nothing to compute against. This is why the machinery is built backwards from what you actually want to know.
Why the null can never be proven
You can compute how likely your data are given the null. You cannot run that backwards to get how likely the null is given your data — that inversion needs a prior, which is the Bayesian route.
In plain terms: the test can tell you “if nothing had changed, results like these would be rare” — never “nothing changed” or “something definitely changed”.
So the two available verdicts are reject and fail to reject. Never accept. A non-significant result means the data didn’t discredit the boring explanation, which is a much weaker statement than the boring explanation being right — see Inconclusive Results.
If you genuinely need to establish that something isn’t worse — a redesign that must not harm conversion — that’s a different design with the hypotheses swapped round: Non-Inferiority Tests.
Two-tailed and one-tailed
- Two-tailed — H₁ is “different in either direction”. You’d act on a significant result whichever way it went
- One-tailed — H₁ is “better”, and you commit in advance to treating a significant loss as no result at all
One-tailed testing halves the p-value for the same data, which is exactly why it’s tempting and almost always wrong. Choosing it after seeing which way the result went is straightforwardly P-Hacking. See One-Tailed vs Two-Tailed Tests.
The honest test: would you ship the variant if it lost significantly? No — then you care about both directions, because a significant loss is information you’d act on. Use two-tailed.
Writing them properly
Most real disputes about test results are disputes about hypotheses that were never written down. Three things to fix before launch:
- Which metric. “Conversion improves” is not a hypothesis until you say conversion of what, over what denominator, in what window — see Metric Design and Overall Evaluation Criterion
- Which population. All visitors, or new visitors, or mobile? The hypothesis names it; discovering it afterwards is post-hoc segmentation
- A mechanism. “Removing the coupon field increases conversion because it stops customers leaving to search for codes” is testable and teaches you something either way. “The new design will perform better” teaches you nothing whichever way it lands
That last point is why Hypothesis Design sits in Experimentation rather than here: the statistics only needs H₀ and H₁: the mechanism is what makes the result worth having.