Tags: statistics concept
Other Statistical Tests
Date: 2026-08-17
The tests you reach for when the outcome isn’t a conversion rate. Proportion Tests cover the large majority of commercial testing; this is the reference for the rest — means, skewed distributions, more than two groups, and paired measurements.
Choosing one
what is the outcome?
│
├─ yes/no per unit (converted, clicked, returned)
│ → PROPORTION TEST — Proportion Tests
│
├─ a number per unit (order value, time, items)
│ │
│ ├─ roughly symmetric, no extreme tail
│ │ → WELCH'S t-TEST
│ │
│ └─ skewed or heavy-tailed ← nearly all money and time data
│ → cap first, then t-test — Winsorisation and Capping
│ → or MANN-WHITNEY, if you accept it answers
│ a different question
│ → or BOOTSTRAP, which is usually the honest answer
│
├─ a number, same units measured twice (before/after on the same people)
│ → PAIRED t-TEST
│
├─ more than two groups
│ → ANOVA, then post-hoc pairwise with a correction
│
└─ a ratio whose numerator and denominator both vary
→ NOT a standard test. Ratio Metrics and the delta method,
or Bootstrapping
Capping before a means test isn’t optional for money data — Winsorisation and Capping — and a ratio of two varying quantities needs its own treatment entirely — Ratio Metrics.
Welch’s t-test
Compares two means without assuming the groups have equal variance. Use Welch’s by default rather than Student’s — it costs almost nothing when variances are equal and is substantially more reliable when they aren’t, which they usually aren’t in commerce.
average order value, capped at the 99th percentile
n mean sd
control 4,120 £42.80 £31.20
variant 4,095 £45.10 £34.60
SE = √(31.20²/4,120 + 34.60²/4,095)
= √(973.44/4,120 + 1,197.16/4,095)
= √(0.23627 + 0.29234)
= √0.52861
= £0.727
t = (45.10 − 42.80) / 0.727 = 2.30 / 0.727 = 3.16
95% CI on the difference = 2.30 ± (1.96 × 0.727)
= £0.87 to £3.73
In plain terms: the variant’s average order is about £2.30 higher, and the data is consistent with the true difference being anywhere from 87p to £3.73. The interval excludes zero, so the difference is unlikely to be chance — but the honest statement is the range, not the £2.30.
Mann-Whitney U
A rank-based test — it throws away the values and keeps only the ordering. Robust to outliers by construction, since the largest value is just “the highest rank”.
The catch that matters: it does not test whether the means differ. It tests whether a randomly chosen value from one group tends to exceed one from the other. Those come apart exactly when the tail matters:
control £20 £20 £20 £20 £20 £20 £20 £20 £20 £20 mean £20
variant £18 £18 £18 £18 £18 £18 £18 £18 £18 £400 mean £56.20
Mann-Whitney: control wins — 9 of 10 variant values are lower
your P&L: variant made nearly 3× the revenue
Never use it as the primary test for a revenue metric. The business cares about the total, and this test is designed to ignore exactly the observations that make up most of it. It’s appropriate for time-on-task, satisfaction scores and similar measures where the ranking is the meaningful thing.
Paired t-test
For the same units measured twice — the same users before and after, or the same pages under two conditions. Test the differences, not the two sets separately.
page before after diff
/product/a 2.41s 2.12s −0.29
/product/b 3.80s 3.31s −0.49
/product/c 1.92s 1.88s −0.04
…
mean diff = −0.31s, sd of diffs = 0.22, n = 40
SE = 0.22 / √40 = 0.0348
t = −0.31 / 0.0348 = −8.9
Pairing removes between-unit variation, which is usually the largest source of noise — so a paired test on 40 pages can be far more powerful than an unpaired test on 400. Same logic as Variance Reduction.
ANOVA
Analysis of variance — tests whether any of three or more group means differ. Its output is a single yes/no about the whole set, which is rarely the question you have.
four landing page variants
ANOVA: F = 4.82, p = 0.0024
→ "at least one differs" ← that's all it says
→ it does NOT say which
then: pairwise comparisons, 6 of them, WITH a correction
— Bonferroni and False Discovery Rate
The correction for those six comparisons is in Bonferroni and False Discovery Rate.
Its value is as a gate: if ANOVA doesn’t reject, stop, and you’ve avoided six comparisons. In practice, most multi-arm commercial tests skip it and go straight to corrected pairwise comparisons against control, which is defensible and simpler — Multivariate Tests.
Chi-squared
Tests whether counts across categories differ from expected. For a 2×2 table it’s algebraically equivalent to a two-proportion z-test — the same question asked twice, so use whichever the tool gives you.
Where it earns its own place is bigger tables — a 2×5 test of whether variant changes the distribution of payment methods chosen, say — and Sample Ratio Mismatch checks, which are a chi-squared goodness-of-fit test against the intended split.
Bootstrapping
Worth naming here as the general escape hatch. When the metric is a ratio, the distribution is ugly, or no formula fits, resampling gives an interval without needing the maths to cooperate. It’s increasingly what tools do underneath — Bootstrapping.
The practical position
- Conversion rate → proportion test. That’s most of what you’ll run
- Revenue → cap, then Welch’s t-test, or bootstrap. Never uncapped, never Mann-Whitney as primary
- More than two arms → correct for multiplicity, whether or not you run ANOVA first
- Check the assumption that actually matters: independence. Every test here assumes observations are independent, and multiple sessions from one user break that regardless of which test you chose — Random Variables, Randomisation Unit
- The test choice is rarely the thing deciding your result. Sample size, metric choice and data quality dominate it — Metric Selection for Tests
Where it interacts
- Proportion Tests — the workhorse this is the complement to
- Skewed and Heavy-Tailed Distributions — the reason money data needs capping before any of these apply
- Bootstrapping — the alternative that avoids most of these choices
- Guide - Choosing a Statistical Test — the decision procedure in full, with the surrounding workflow