Tags: statistics concept

Other Statistical Tests

Date: 2026-08-17


The tests you reach for when the outcome isn’t a conversion rate. Proportion Tests cover the large majority of commercial testing; this is the reference for the rest — means, skewed distributions, more than two groups, and paired measurements.


Choosing one

what is the outcome?
│
├─ yes/no per unit (converted, clicked, returned)
│     → PROPORTION TEST — Proportion Tests
│
├─ a number per unit (order value, time, items)
│     │
│     ├─ roughly symmetric, no extreme tail
│     │     → WELCH'S t-TEST
│     │
│     └─ skewed or heavy-tailed  ← nearly all money and time data
│           → cap first, then t-test — Winsorisation and Capping
│           → or MANN-WHITNEY, if you accept it answers
│             a different question
│           → or BOOTSTRAP, which is usually the honest answer
│
├─ a number, same units measured twice (before/after on the same people)
│     → PAIRED t-TEST
│
├─ more than two groups
│     → ANOVA, then post-hoc pairwise with a correction
│
└─ a ratio whose numerator and denominator both vary
      → NOT a standard test. Ratio Metrics and the delta method,
        or Bootstrapping

Capping before a means test isn’t optional for money data — Winsorisation and Capping — and a ratio of two varying quantities needs its own treatment entirely — Ratio Metrics.

Welch’s t-test

Compares two means without assuming the groups have equal variance. Use Welch’s by default rather than Student’s — it costs almost nothing when variances are equal and is substantially more reliable when they aren’t, which they usually aren’t in commerce.

average order value, capped at the 99th percentile

           n        mean      sd
control  4,120    £42.80    £31.20
variant  4,095    £45.10    £34.60

SE  =  √(31.20²/4,120  +  34.60²/4,095)
    =  √(973.44/4,120  +  1,197.16/4,095)
    =  √(0.23627 + 0.29234)
    =  √0.52861
    =  £0.727

t   =  (45.10 − 42.80) / 0.727  =  2.30 / 0.727  =  3.16

95% CI on the difference  =  2.30 ± (1.96 × 0.727)
                          =  £0.87 to £3.73

In plain terms: the variant’s average order is about £2.30 higher, and the data is consistent with the true difference being anywhere from 87p to £3.73. The interval excludes zero, so the difference is unlikely to be chance — but the honest statement is the range, not the £2.30.

Mann-Whitney U

A rank-based test — it throws away the values and keeps only the ordering. Robust to outliers by construction, since the largest value is just “the highest rank”.

The catch that matters: it does not test whether the means differ. It tests whether a randomly chosen value from one group tends to exceed one from the other. Those come apart exactly when the tail matters:

control  £20 £20 £20 £20 £20 £20 £20 £20 £20 £20     mean £20
variant  £18 £18 £18 £18 £18 £18 £18 £18 £18 £400    mean £56.20

Mann-Whitney: control wins — 9 of 10 variant values are lower
your P&L:     variant made nearly 3× the revenue

Never use it as the primary test for a revenue metric. The business cares about the total, and this test is designed to ignore exactly the observations that make up most of it. It’s appropriate for time-on-task, satisfaction scores and similar measures where the ranking is the meaningful thing.

Paired t-test

For the same units measured twice — the same users before and after, or the same pages under two conditions. Test the differences, not the two sets separately.

page          before   after   diff
/product/a     2.41s   2.12s   −0.29
/product/b     3.80s   3.31s   −0.49
/product/c     1.92s   1.88s   −0.04
…

mean diff = −0.31s, sd of diffs = 0.22, n = 40
SE = 0.22 / √40 = 0.0348
t  = −0.31 / 0.0348 = −8.9

Pairing removes between-unit variation, which is usually the largest source of noise — so a paired test on 40 pages can be far more powerful than an unpaired test on 400. Same logic as Variance Reduction.

ANOVA

Analysis of variance — tests whether any of three or more group means differ. Its output is a single yes/no about the whole set, which is rarely the question you have.

four landing page variants

ANOVA:  F = 4.82,  p = 0.0024
        → "at least one differs"  ← that's all it says
        → it does NOT say which

then: pairwise comparisons, 6 of them, WITH a correction
      — Bonferroni and False Discovery Rate

The correction for those six comparisons is in Bonferroni and False Discovery Rate.

Its value is as a gate: if ANOVA doesn’t reject, stop, and you’ve avoided six comparisons. In practice, most multi-arm commercial tests skip it and go straight to corrected pairwise comparisons against control, which is defensible and simpler — Multivariate Tests.

Chi-squared

Tests whether counts across categories differ from expected. For a 2×2 table it’s algebraically equivalent to a two-proportion z-test — the same question asked twice, so use whichever the tool gives you.

Where it earns its own place is bigger tables — a 2×5 test of whether variant changes the distribution of payment methods chosen, say — and Sample Ratio Mismatch checks, which are a chi-squared goodness-of-fit test against the intended split.

Bootstrapping

Worth naming here as the general escape hatch. When the metric is a ratio, the distribution is ugly, or no formula fits, resampling gives an interval without needing the maths to cooperate. It’s increasingly what tools do underneath — Bootstrapping.

The practical position

  • Conversion rate → proportion test. That’s most of what you’ll run
  • Revenue → cap, then Welch’s t-test, or bootstrap. Never uncapped, never Mann-Whitney as primary
  • More than two arms → correct for multiplicity, whether or not you run ANOVA first
  • Check the assumption that actually matters: independence. Every test here assumes observations are independent, and multiple sessions from one user break that regardless of which test you chose — Random Variables, Randomisation Unit
  • The test choice is rarely the thing deciding your result. Sample size, metric choice and data quality dominate it — Metric Selection for Tests

Where it interacts