Tags: experimentation map
Experimentation MOC
Date: 2026-08-16
Experimentation is the only method that establishes what a change caused rather than what accompanied it. Everything here is either protecting that property or exploiting it.
Ordered by the life of a test: design, run, read, then the programme around it. Statistics owns the maths — power, p-values, corrections. This domain owns the practice.
Ground
- Controlled Experiments — the parent form: two groups differing in one thing, chosen at random. An A/B test is one instance of it, a geo holdout is another
- Why Randomisation Works — how random assignment removes confounders you never thought to measure
- Randomisation Unit — user, session, device, cookie, account. The choice that determines what you can measure and what leaks
- Assignment and Bucketing — hashing an identifier into a variant deterministically, so the same user sees the same thing
- A-A Tests — running the same experience against itself to prove the machinery is honest before trusting it
- Experiment Assignment Tracking — recording who saw what, without which no analysis is possible
Designing a test
Most experiments are won or lost here, before any traffic arrives.
- Hypothesis Design — a testable statement with a mechanism, not a preference with a metric attached
- Overall Evaluation Criterion — the single primary metric decided in advance, and why more than one is a licence to pick a winner afterwards
- Guardrail Metrics — the things that must not get worse, whatever the primary metric does
- Secondary and Diagnostic Metrics — explaining why a result happened without letting them decide the result
- Metric Selection for Tests — sensitivity versus meaningfulness, and why the most important metric is often the wrong one to test on
- Test Duration — business cycles, not just sample size; why a week is the minimum unit
- Traffic Allocation — even splits, unequal splits, ramping, and what each costs in power
- Pre-Registration — writing the analysis plan before seeing data, which is what makes the p-value mean anything
- Test Prioritisation — ICE, PIE and their honest limitations as scoring theatre
Running a test
- Experiment QA — the checks before traffic: both variants render, tracking fires, the split is right, nothing leaks
- Client-Side vs Server-Side Testing — where the variant decision is made, and everything that follows from it
- Flicker and Flash of Original Content — the client-side tax: the original rendering before the variant applies, and what it does to results
- Interaction Effects — concurrent tests affecting each other, and when that actually matters
- Novelty and Primacy Effects — regular users reacting to change itself, which fades
- Seasonality in Tests — running across a peak, a payday or a campaign and mistaking it for treatment
- Sample Pollution — bots, internal traffic and duplicate assignment contaminating both arms
- Stopping Rules — deciding in advance what ends a test, so the decision isn’t made by whoever’s watching
Test types
- A-B Tests — a controlled experiment run on users, on your own site: the conventions, the vocabulary, and the practical shape of one
- Multivariate Tests — several elements at once, and the traffic cost that usually rules them out
- Split URL Tests — separate pages rather than modified ones, and the SEO and redirect implications
- Multi-Armed Bandits — allocating traffic towards the winner during the test, and what you give up to get it
- Holdout Groups — a permanently untreated slice, to measure cumulative effect over a year rather than per-test
- Switchback Tests — alternating treatment over time when users can’t be independently assigned
- Painted Door Tests — measuring demand for something that doesn’t exist yet, and the ethics of it
- Rollouts as Experiments — when a staged rollout supports a causal claim and when it is only a release
- Personalisation Tests — testing targeting rules rather than experiences, and why they need different analysis
Reading results
- Reading a Test Result — the ordered checks before believing anything: SRM, guardrails, duration, primary metric, then everything else
- Segmentation (test results) — slicing a result post hoc, and why almost every finding from it is noise
- Heterogeneous Treatment Effects — a real difference between groups, and how to distinguish it from the above
- Winner’s Curse — why the winners you ship consistently underdeliver against their test result
- Inconclusive Results — the most common outcome, what it does and doesn’t tell you, and why it isn’t failure
- Non-Inferiority Tests — proving something isn’t worse, which is a different design from proving it’s better
- Post-Test Validation — checking the shipped change delivered what the test promised
The programme
Individual tests matter less than the system that produces them.
- Experimentation Velocity — tests per period as the real driver of cumulative gain
- Win Rate and Expected Value — most tests fail; the programme’s value is the arithmetic across all of them
- Experiment Archive — the searchable record of what was tried, which is the asset that outlives every individual test
- Institutional Learning — turning results into beliefs, and revisiting beliefs when results contradict them
- Experimentation Maturity — the stages an organisation moves through, and the bottleneck at each
- The HiPPO Problem — the highest-paid person’s opinion deciding what ships, and testing as the way to move the argument off rank
- Ethics of Experimentation — consent, harm, and the things that shouldn’t be tested on people
- Testing and Compliance — regulated claims, pricing tests, and where experimentation meets legal constraint in the UK
Borders
Filed elsewhere, needed constantly here.
- Statistics — Sample Ratio Mismatch · Statistical Power · Minimum Detectable Effect · Sample Size Calculation · P-Values · Confidence Intervals · Peeking · The Multiple Comparisons Problem · Variance Reduction · Bayesian vs Frequentist
- Analytics — Event Taxonomy Design · Sessionisation · Identity Stitching · Metric Design · Tracking Plans
- Web Development — Feature Flags · Progressive Delivery · Rendering Strategies
- User Experience — Usability Testing · Qualitative vs Quantitative Research
- Commerce & Growth — Incrementality Testing · Geo Holdout Tests · Price Testing