Tags: experimentation map

Experimentation MOC

Date: 2026-08-16


Experimentation is the only method that establishes what a change caused rather than what accompanied it. Everything here is either protecting that property or exploiting it.


Ordered by the life of a test: design, run, read, then the programme around it. Statistics owns the maths — power, p-values, corrections. This domain owns the practice.

Ground

  • Controlled Experiments — the parent form: two groups differing in one thing, chosen at random. An A/B test is one instance of it, a geo holdout is another
  • Why Randomisation Works — how random assignment removes confounders you never thought to measure
  • Randomisation Unit — user, session, device, cookie, account. The choice that determines what you can measure and what leaks
  • Assignment and Bucketing — hashing an identifier into a variant deterministically, so the same user sees the same thing
  • A-A Tests — running the same experience against itself to prove the machinery is honest before trusting it
  • Experiment Assignment Tracking — recording who saw what, without which no analysis is possible

Designing a test

Most experiments are won or lost here, before any traffic arrives.

  • Hypothesis Design — a testable statement with a mechanism, not a preference with a metric attached
  • Overall Evaluation Criterion — the single primary metric decided in advance, and why more than one is a licence to pick a winner afterwards
  • Guardrail Metrics — the things that must not get worse, whatever the primary metric does
  • Secondary and Diagnostic Metrics — explaining why a result happened without letting them decide the result
  • Metric Selection for Tests — sensitivity versus meaningfulness, and why the most important metric is often the wrong one to test on
  • Test Duration — business cycles, not just sample size; why a week is the minimum unit
  • Traffic Allocation — even splits, unequal splits, ramping, and what each costs in power
  • Pre-Registration — writing the analysis plan before seeing data, which is what makes the p-value mean anything
  • Test Prioritisation — ICE, PIE and their honest limitations as scoring theatre

Running a test

Test types

  • A-B Tests — a controlled experiment run on users, on your own site: the conventions, the vocabulary, and the practical shape of one
  • Multivariate Tests — several elements at once, and the traffic cost that usually rules them out
  • Split URL Tests — separate pages rather than modified ones, and the SEO and redirect implications
  • Multi-Armed Bandits — allocating traffic towards the winner during the test, and what you give up to get it
  • Holdout Groups — a permanently untreated slice, to measure cumulative effect over a year rather than per-test
  • Switchback Tests — alternating treatment over time when users can’t be independently assigned
  • Painted Door Tests — measuring demand for something that doesn’t exist yet, and the ethics of it
  • Rollouts as Experiments — when a staged rollout supports a causal claim and when it is only a release
  • Personalisation Tests — testing targeting rules rather than experiences, and why they need different analysis

Reading results

The programme

Individual tests matter less than the system that produces them.

Borders

Filed elsewhere, needed constantly here.