Tags: statistics concept

Hypothesis Testing

Date: 2026-08-16


Assume nothing happened, then ask how surprising your data would be if that were true. Surprising enough, and you reject the assumption. You never prove the thing you wanted to show — you only discredit its opposite.


What it is

Hypothesis testing is a procedure for deciding whether observed data is surprising enough to discredit a stated assumption — almost always the assumption that nothing is happening.

The frame

It’s proof by contradiction, run on probabilities.

  1. Assume there is no effect. This is the null hypothesis, written H₀
  2. Work out what data would look like if that were true — not one outcome, but the whole distribution of outcomes you’d see across repeated experiments
  3. Compare your actual data to that distribution. How far into the tail is it?
  4. If it’s far enough out, reject the assumption. The threshold for “far enough” is Statistical Significance

The output is a p-value: the probability of seeing data at least this extreme if the null were true.

Worked

Control converts 3.00% on 53,000 visitors; variant converts 3.09% on 53,000.

Assume no real difference. Under that assumption, both are samples from one population with a true rate around 3.045%. The observed difference between two such samples has a standard error of:

SE = √[ 2 × p̄(1−p̄) / n ]
   = √[ 2 × 0.03045 × 0.96955 / 53,000 ]
   = √[ 0.05904 / 53,000 ]
   = √0.000001114
   = 0.001056  →  0.1056 percentage points

The observed difference is 3.09% − 3.00% = 0.09 percentage points.

z = 0.09 ÷ 0.1056 = 0.85

A difference 0.85 standard errors from zero. Under the null, differences that size or larger happen about 40% of the time. That is not surprising at all, so the null survives.

In plain terms: if the two versions were genuinely identical, you’d still see a gap this big or bigger in roughly two experiments out of five, just from who happened to turn up. There’s nothing here to act on.

What it can and can’t conclude

The asymmetry is the whole thing, and it trips everyone:

You can sayYou cannot say
”The data are unlikely under the null, so I reject it""The null is false"
"The data are consistent with the null""There is no effect"
"I detected an effect""I proved the effect exists”

Failing to reject is not accepting. A test that comes back non-significant is consistent with no effect, with a small effect, and with a large effect you were too underpowered to see — see Statistical Power and Inconclusive Results. The three are not distinguished by the p-value; look at the confidence interval to tell them apart.

Why this is counterintuitive

The machinery answers a question nobody asks. You want “how likely is it that my variant is better?” Hypothesis testing answers “how likely is data like this, assuming my variant is identical?”

Those are different questions and the answers are not interchangeable — inverting them is the Base Rate Fallacy. The framework that does answer the question you asked is Bayesian, at the cost of requiring a prior.

Where it goes wrong in practice

Each of these breaks the “if the null were true” calculation, which means the p-value stops describing anything.