Tags: statistics map

Statistics MOC

Date: 2026-08-16


Statistics is the discipline of saying how much you know, given that you only measured some of it. Everything here is a way of quantifying that gap — or a way it goes wrong.


Strictly dependency-ordered, and the one domain where reading out of order doesn’t work. Inference assumes sampling; sequential testing assumes multiplicity; causal inference assumes everything above it.

Working from a real test rather than a topic? Guide - Statistics for CRO runs one end to end with the arithmetic done, and links back here for every concept.

Ground

The vocabulary everything else is written in.

Sampling

The bridge from what you measured to what you’re claiming.

  • Sampling Methods — random, stratified, cluster, convenience, and what each buys
  • Sampling Error — the unavoidable gap between a sample and its population, its shrinking rate, and why small samples resemble the population far less than intuition insists
  • Standard Error — how much a sample statistic would bounce around if you did it again
  • Selection Bias — the sample being systematically unlike the population, which no sample size fixes
  • Survivorship Bias — measuring only the things still present to measure
  • Sample Ratio Mismatch — split that doesn’t match the intended allocation; a signal that assignment itself is broken

Inference

The core machinery. Slow down here.

Tests

Two notes, because in practice you run one of them.

  • Proportion Tests — comparing rates directly. The workhorse: nearly every conversion test is this, and the chi-squared and z-test forms are the same question asked twice
  • Other Statistical Tests — t-tests for means, Mann-Whitney when the distribution won’t cooperate, ANOVA for more than two groups. Reference depth, reached for rarely
  • Bootstrapping — building a confidence interval by resampling instead of by formula. What the tools increasingly do underneath, and the honest answer for ratio metrics

Multiplicity and stopping

Where honest analysts produce false results without noticing.

Bayesian

A different frame for the same questions. Worth understanding even if you never run one.

Relationships

Causal inference

Establishing that something caused something, which observation alone never does.

Traps

Each one has produced a confident, wrong, published result. They’re listed separately because they’re the recurring diagnoses.

Applied

Where the maths meets the metrics in this vault.

  • Ratio Metrics — why the variance of a ratio isn’t the variance of its parts, and what that does to significance
  • Variance Reduction — CUPED and covariates: using what you knew before the test to get the same certainty from less traffic
  • Winsorisation and Capping — capping extreme order values so one £5,000 purchase can’t decide a test. Winsorise replaces, trim deletes, and the cap must be chosen before you look
  • Metric Sensitivity — why revenue per visitor needs vastly more traffic than conversion rate to move detectably
  • Communicating Uncertainty — presenting a range to people who want a number

Borders

Filed elsewhere, needed constantly here.