Tags: statistics concept

Standard Error

Date: 2026-08-16


How much your result would bounce around if you ran the same measurement again. It’s the standard deviation of a statistic rather than of the data — and confusing the two is the most common misreading in the field.


What it is

The standard error is the standard deviation of a sample statistic across hypothetical repeated samples.

For a proportion, substituting the binomial variance:

Worked, on a 3% conversion rate with 53,000 visitors:

p(1−p)        = 0.03 × 0.97        = 0.0291
÷ n           = 0.0291 ÷ 53,000    = 0.000000549
√             = 0.000741           = 0.0741 percentage points

So a measured 3.00% carries a standard error of 0.074pp. Roughly two standard errors either side gives the 95% interval: 2.85% to 3.15%.

Standard deviation versus standard error

The distinction, stated plainly because it’s the one that trips people:

DescribesBehaviour as n grows
Standard deviationHow spread out the data isStays the same — it’s a property of the population
Standard errorHow precisely you’ve measured a statisticShrinks, as √n

In plain terms: your customers don’t become more similar to each other as you collect more data. Your estimate of their average becomes more precise. The first is standard deviation, the second is standard error, and only the second improves with sample size.

The square root is the whole economics of testing

Precision improves with the square root of sample size, which means:

to halve the standard error  →  4× the data
to third it                  →  9× the data

Run this backwards and you get the relationship that dominates test planning: sample size scales with the square of the effect you want to detect. Halving your MDE quadruples the traffic — Minimum Detectable Effect.

It’s also why “run it a bit longer” is weak. Doubling duration improves precision by about 30%, not 100%.

The standard error of a difference

What a test actually uses — comparing two arms, not describing one:

Variances add. The standard error of a difference is larger than either arm’s alone, which is why comparing two measurements is harder than measuring one.

Worked, both arms at 53,000, pooled around 3.045%:

2 × 0.03045 × 0.96955 ÷ 53,000  = 0.000001114
√                                = 0.001056  →  0.106pp

That 0.106pp is the number every conversion test’s significance and interval is computed from — see Hypothesis Testing and Confidence Intervals.

Reading it

  • A wide standard error means an imprecise measurement, whatever the point estimate says. The estimate isn’t wrong, it’s just uninformative
  • It scales with √n, so small segments are noisy fast. Splitting a test into four segments quarters n and doubles the standard error in each — which is most of why post-hoc segments produce false findings
  • It assumes independence. Counting the same user twice understates it, which overstates significance — Randomisation Unit
  • It assumes the sampling distribution is roughly normal, which for heavy-tailed metrics at modest samples it isn’t — The Central Limit Theorem, Bootstrapping

Where it feeds

Almost everything: Confidence Intervals (± 1.96 × SE), P-Values (the test statistic is the difference divided by SE), Sample Size Calculation (solving for the n that makes SE small enough), and Sampling Error (the conceptual gap the SE quantifies).