Tags: statistics concept
Standard Error
Date: 2026-08-16
How much your result would bounce around if you ran the same measurement again. It’s the standard deviation of a statistic rather than of the data — and confusing the two is the most common misreading in the field.
What it is
The standard error is the standard deviation of a sample statistic across hypothetical repeated samples.
For a proportion, substituting the binomial variance:
Worked, on a 3% conversion rate with 53,000 visitors:
p(1−p) = 0.03 × 0.97 = 0.0291
÷ n = 0.0291 ÷ 53,000 = 0.000000549
√ = 0.000741 = 0.0741 percentage points
So a measured 3.00% carries a standard error of 0.074pp. Roughly two standard errors either side gives the 95% interval: 2.85% to 3.15%.
Standard deviation versus standard error
The distinction, stated plainly because it’s the one that trips people:
| Describes | Behaviour as n grows | |
|---|---|---|
| Standard deviation | How spread out the data is | Stays the same — it’s a property of the population |
| Standard error | How precisely you’ve measured a statistic | Shrinks, as √n |
In plain terms: your customers don’t become more similar to each other as you collect more data. Your estimate of their average becomes more precise. The first is standard deviation, the second is standard error, and only the second improves with sample size.
The square root is the whole economics of testing
Precision improves with the square root of sample size, which means:
to halve the standard error → 4× the data
to third it → 9× the data
Run this backwards and you get the relationship that dominates test planning: sample size scales with the square of the effect you want to detect. Halving your MDE quadruples the traffic — Minimum Detectable Effect.
It’s also why “run it a bit longer” is weak. Doubling duration improves precision by about 30%, not 100%.
The standard error of a difference
What a test actually uses — comparing two arms, not describing one:
Variances add. The standard error of a difference is larger than either arm’s alone, which is why comparing two measurements is harder than measuring one.
Worked, both arms at 53,000, pooled around 3.045%:
2 × 0.03045 × 0.96955 ÷ 53,000 = 0.000001114
√ = 0.001056 → 0.106pp
That 0.106pp is the number every conversion test’s significance and interval is computed from — see Hypothesis Testing and Confidence Intervals.
Reading it
- A wide standard error means an imprecise measurement, whatever the point estimate says. The estimate isn’t wrong, it’s just uninformative
- It scales with √n, so small segments are noisy fast. Splitting a test into four segments quarters n and doubles the standard error in each — which is most of why post-hoc segments produce false findings
- It assumes independence. Counting the same user twice understates it, which overstates significance — Randomisation Unit
- It assumes the sampling distribution is roughly normal, which for heavy-tailed metrics at modest samples it isn’t — The Central Limit Theorem, Bootstrapping
Where it feeds
Almost everything: Confidence Intervals (± 1.96 × SE), P-Values (the test statistic is the difference divided by SE), Sample Size Calculation (solving for the n that makes SE small enough), and Sampling Error (the conceptual gap the SE quantifies).