Tags: statistics map
Statistics MOC
Date: 2026-08-16
Statistics is the discipline of saying how much you know, given that you only measured some of it. Everything here is a way of quantifying that gap — or a way it goes wrong.
Strictly dependency-ordered, and the one domain where reading out of order doesn’t work. Inference assumes sampling; sequential testing assumes multiplicity; causal inference assumes everything above it.
Working from a real test rather than a topic? Guide - Statistics for CRO runs one end to end with the arithmetic done, and links back here for every concept.
Ground
The vocabulary everything else is written in.
- Descriptive vs Inferential Statistics — summarising what you measured versus claiming something about what you didn’t
- Populations and Samples — the distinction the whole field rests on, and why “the population” is usually hypothetical
- Random Variables — a number whose value depends on a process you can describe but not predict
- Probability Distributions — the shape of how often each value occurs, and why the shape decides the method
- The Normal Distribution — the one everything defaults to, its conditions, and what breaks when they don’t hold
- Binomial and Bernoulli Distributions — the maths of yes/no outcomes, which is every conversion rate
- Skewed and Heavy-Tailed Distributions — revenue, session duration, order value: the ones where the mean misleads
- Mean Median and Mode — three answers to “typical”, and when each lies
- Variance and Standard Deviation — spread, and why it matters more than the average for decisions
- Percentiles and Quantiles — reading a distribution by position rather than by average
- The Central Limit Theorem — why sample averages go normal even when the underlying data doesn’t, and the conditions on that
Sampling
The bridge from what you measured to what you’re claiming.
- Sampling Methods — random, stratified, cluster, convenience, and what each buys
- Sampling Error — the unavoidable gap between a sample and its population, its shrinking rate, and why small samples resemble the population far less than intuition insists
- Standard Error — how much a sample statistic would bounce around if you did it again
- Selection Bias — the sample being systematically unlike the population, which no sample size fixes
- Survivorship Bias — measuring only the things still present to measure
- Sample Ratio Mismatch — split that doesn’t match the intended allocation; a signal that assignment itself is broken
Inference
The core machinery. Slow down here.
- Hypothesis Testing — the whole frame: assume nothing happened, ask how surprising the data would be
- Null and Alternative Hypotheses — what you’re actually testing, and why the null is the one you can reject but never prove
- P-Values — the most misread number in the field; what it is, and the four things it isn’t
- Statistical Significance — a threshold decision, not a discovery, and the arbitrariness of 0.05
- Type I and Type II Errors — false positive and false negative, and the trade you’re always making between them
- Statistical Power — the chance of detecting a real effect; the number most people skip and then regret
- Minimum Detectable Effect — the smallest effect a test can find, and why deciding it up front changes the design
- Sample Size Calculation — assembling baseline, MDE, power and significance into one number
- Confidence Intervals — the range, why it’s more honest than a point estimate, and what “95% confident” doesn’t mean
- One-Tailed vs Two-Tailed Tests — the choice that halves your p-value and is almost always wrong to make
- Effect Size — how big, as distinct from how certain. The question significance never answers
- Practical vs Statistical Significance — a real effect too small to be worth shipping, which is how a testing programme wastes a year while passing every check
Tests
Two notes, because in practice you run one of them.
- Proportion Tests — comparing rates directly. The workhorse: nearly every conversion test is this, and the chi-squared and z-test forms are the same question asked twice
- Other Statistical Tests — t-tests for means, Mann-Whitney when the distribution won’t cooperate, ANOVA for more than two groups. Reference depth, reached for rarely
- Bootstrapping — building a confidence interval by resampling instead of by formula. What the tools increasingly do underneath, and the honest answer for ratio metrics
Multiplicity and stopping
Where honest analysts produce false results without noticing.
- The Multiple Comparisons Problem — testing twenty things at 95% confidence and expecting one false positive by construction
- Bonferroni and False Discovery Rate — the two families of correction, and what each costs in power
- Peeking — checking a test in flight, and how it inflates the false positive rate far past the nominal one
- Sequential Testing — designs built to be checked repeatedly, and the price paid for that licence
- Always-Valid Inference — confidence intervals that stay honest under continuous monitoring
- P-Hacking — the deliberate version
- The Garden of Forking Paths — the accidental version, which is far more common and much harder to see
Bayesian
A different frame for the same questions. Worth understanding even if you never run one.
- Bayesian vs Frequentist — what each is actually claiming, and why they answer different questions
- Prior Likelihood and Posterior — the mechanism, in one worked update
- Credible Intervals — the interval that means what people wrongly think a confidence interval means
- Probability to Beat Control — the output most Bayesian test tools report, and how to read it honestly
- Choosing a Prior — where the subjectivity lives, and why “uninformative” isn’t neutral
Relationships
- Correlation — measuring co-movement, and what a coefficient does and doesn’t tell you
- Correlation and Causation — the specific reasons the leap fails, beyond the slogan
- Confounding Variables — the third thing causing both, and why it’s the default explanation
- Linear Regression — fitting a line, reading the coefficients, checking the assumptions
- Multiple Regression — controlling for things, and the illusion of control it creates
- Logistic Regression — regression for binary outcomes, which is most outcomes here
- Regression to the Mean — extremes moving towards average on their own, and the interventions it falsely credits
Causal inference
Establishing that something caused something, which observation alone never does.
- Randomised Controlled Trials — why randomisation is the only thing that removes confounders you haven’t thought of
- Counterfactuals — the unobservable comparison every causal claim is really making
- Difference-in-Differences — using a comparison group’s trend as the counterfactual
- Propensity Score Matching — constructing a comparable group from observational data, and its limits
- Natural Experiments — randomisation the world did for you
Traps
Each one has produced a confident, wrong, published result. They’re listed separately because they’re the recurring diagnoses.
- Simpson’s Paradox — every segment moving one way while the total moves the other
- Base Rate Fallacy — ignoring how rare something is when reading a test that detects it
- Outliers and Robust Statistics — spotting the values distorting a result, and the estimators that resist them
Applied
Where the maths meets the metrics in this vault.
- Ratio Metrics — why the variance of a ratio isn’t the variance of its parts, and what that does to significance
- Variance Reduction — CUPED and covariates: using what you knew before the test to get the same certainty from less traffic
- Winsorisation and Capping — capping extreme order values so one £5,000 purchase can’t decide a test. Winsorise replaces, trim deletes, and the cap must be chosen before you look
- Metric Sensitivity — why revenue per visitor needs vastly more traffic than conversion rate to move detectably
- Communicating Uncertainty — presenting a range to people who want a number
Borders
Filed elsewhere, needed constantly here.
- Experimentation — Randomisation Unit · Guardrail Metrics · A-A Tests · Reading a Test Result
- Analytics — Sessionisation · Data Sampling · Metric Design · Segmentation (analysis)
- Commerce & Growth — Incrementality Testing · Geo Holdout Tests