Tags: experimentation guide
Guide - Running an Experiment
Date: 2026-09-28
One test, from an idea someone wants to test to an entry in the archive, as a runbook. The statistics are rarely what goes wrong. The decisions made while the test runs are: when to stop, what counts as a win, whether a bug means starting again, what to tell the person asking for early results. Every one of those is decided in writing before launch, so nobody has to make it with the data in front of them.
Running an experiment is the operational procedure around a controlled test: the documents, sign-offs, checks and pre-agreed decisions that make its result trustworthy and its outcome actionable.
Where this sits among the guides:
- Guide - Statistics for CRO — the arithmetic: sample size, reading the interval, pounds
- Guide - Ecommerce CRO §10–15 — the same lifecycle with ecommerce depth: exposure, underpowered sites, the contribution/returns/cohort checks on reading a result
- This guide — the runbook for any site: the five documents a test produces, what’s agreed before launch, and what to do when something goes wrong mid-flight
Running example: finding F3 from Guide - Running a Conversion Audit — delivery cost first seen at checkout, graded B, sent to TEST. Same retailer: 120,000 sessions a month, 72,000 of which view a product, 4.0% of product-viewing sessions ending in an order, £23.01 contribution per order.
The lifecycle
STAGE DOCUMENT SIGNED OFF BY GATE TO THE NEXT STAGE
1 Plan the test plan owner + analyst plan frozen, sample size fits the calendar
2 Build & QA the launch check developer + analyst every launch-check line ticked
3 Launch — owner ramp checks clean at 48h
4 Run the daily log analyst planned end date reached (or a harm stop)
5 Decide the decision record owner + approver the pre-agreed rule applied, not re-argued
6 Close the archive entry analyst code removed, archive written, re-read booked
Four roles, even if two are the same person: the owner (whose idea it is and who wants it to win — which is why they don’t read the result alone), the analyst, the developer, and the approver who can ship. Naming them matters most when the approver is senior — The HiPPO Problem.
1. Plan — the document that makes the result mean something
Everything the analysis depends on is written before any data exists. A result read against a plan written afterwards measures nothing, because every choice — metric, segment, duration — can be fitted to the data. Pre-Registration owns the argument.
The test plan, filled in for F3:
TEST PLAN status: FROZEN 2026-10-05
ID / name EXP-041 delivery cost on product page
Source audit F3 (grade B): 9% of checkout starters exit at the delivery
step vs 4% of those who'd seen the delivery page; "how much is
delivery" is a top-3 support ticket
Hypothesis BECAUSE delivery cost is first shown at checkout
WE BELIEVE the surprise at the point of commitment causes exits
SO IF WE show the delivery cost and free-delivery threshold under
the price on every product page
WE EXPECT more product-viewing sessions to end in an order
Unit user (hashed first-party ID); analysed per exposed session
Exposure sessions that view a product page — assigned on first view
Split 50 / 50 after a 48h ramp at 10%
Primary order conversion of exposed sessions baseline 4.0%
Guardrails contribution per exposed session (−2% = harm)
add-to-basket rate (watch: honest cost may deter early)
product page LCP p75 (+200ms = harm)
JS error rate
Diagnostic delivery-step exit rate among checkout starters
(tests the MECHANISM, never decides the result)
MDE / power 10% relative, 80% power, α 0.05 two-sided
Sample 39,431 exposed sessions per arm
Duration 5 whole weeks (4.8 weeks at 72,000 exposures a month)
Calendar avoids the November sale; no bank holiday weekend in window
Decision rules see below — agreed before launch
Harm stop primary down > 5% relative with p < 0.01 at a weekly check,
or any guardrail at its harm line two checks running
Segments mobile / desktop only, named here; anything else is exploratory
Collisions none on PDP; checkout test EXP-039 ends before this starts
The parts of it that carry the most weight, in order:
The primary metric is measured on everyone exposed, not on people who reach checkout. The obvious metric — checkout completion — is where the mechanism acts, but the change is on the product page, so it also changes who reaches checkout. Showing delivery cost early may turn away some people at the product page who would have abandoned at checkout anyway; checkout completion then rises with no one extra buying. Conditioning on anything that happens after assignment compares two differently selected groups — Selection Bias. The diagnostic metric is conditioned that way, which is exactly why it only explains and never decides — Secondary and Diagnostic Metrics.
The MDE is the audit’s upper bound, and the plan says so. The sample size, worked:
baseline p₁ 0.040
target p₂ (10% relative) 0.044
difference d 0.004
pooled p̄ 0.042
2 × p̄ × (1 − p̄) 0.080472
n per arm = 7.84 × 0.080472 ÷ 0.004²
= 0.630900 ÷ 0.000016 = 39,431
total 78,862 ÷ 72,000 a month = 1.10 months = 4.8 weeks → 5 whole weeks
7.84 is (1.96 + 0.84)², the constants for 95% significance and 80% power — Sample Size Calculation. A 5% MDE would need 154,130 per arm: 4.3 months, which the calendar won’t hold.
The catch: the audit sized F3’s upper bound at about a 9.6% lift, and this test can only reliably detect 10%. So the plan is honest that it’s powered to catch F3 only if F3 is about as good as it could possibly be. A realistic effect of 3–5% will most likely come back inconclusive. That goes in the plan, and the decision rules say what inconclusive means, rather than discovering it in week five — Minimum Detectable Effect.
In plain terms: this test can hear a shout but not a normal voice. If it hears nothing, that isn’t evidence of silence.
The decision rules, the part most plans leave out:
IF THEN
SRM at any check invalid — find the cause, fix, restart
primary up, significant; guardrails clear; SHIP
diagnostic moved as predicted
primary up, significant; diagnostic didn't move SHIP, record the mechanism as NOT supported
primary up, significant; a guardrail at harm DON'T SHIP — redesign for the guardrail
interval spans zero, excludes harm beyond −2% SHIP as non-inferior — we'd show delivery
cost anyway; record "no detectable effect"
interval includes harm beyond −2% DON'T SHIP — re-run only with more power
harm stop triggered stop, roll back, write it up as a loss
The fourth row from the bottom is the one that saves the argument in week five. Showing delivery cost up front is something the business would do regardless — it’s clearer, and UK pricing-transparency rules push the same way [CHECK: whether DMCC Act 2024 drip-pricing rules treat delivery charges as mandatory fees for headline pricing]. So the real question is “does it hurt?”, and an inconclusive result answers that — Non-Inferiority Tests · Inconclusive Results. If you wouldn’t ship regardless, this row would say “don’t ship” instead. Decide which before launch.
Also in planning: Overall Evaluation Criterion for choosing one primary, Guardrail Metrics for choosing the harm lines, Test Duration for the calendar, Randomisation Unit for the unit.
2. Build and QA — the launch check
Experiment QA owns the procedure; Guide - Ecommerce CRO §12 has the ecommerce list. The document is short on purpose, so every line actually gets ticked:
LAUNCH CHECK EXP-041 ☐ = not yet
☐ both arms render — each device class and browser over 2% of traffic
☐ variant copy correct for every delivery rule (threshold, oversize, international)
☐ full purchase completes in both arms on staging; one live order in production
☐ assignment fires once, is stable across page loads and sessions, reaches analytics
☐ exposure logged on first product view, not on site entry
☐ bots excluded from assignment
☐ no flicker — variant present in first render
☐ LCP p75 unchanged in a lab comparison of the two arms
☐ variant survives the CDN cache; cache key includes the arm
☐ consent: declined users are either assigned and unmeasured in BOTH arms, or neither
☐ no other test or personalisation rule touches the product page price block
☐ kill switch tested — the flag turns the variant off without a deploy
☐ plan frozen and linked
The kill-switch line is the one that gets skipped, and it’s the one you need at 18:00 on a Friday. Serve the variant behind a flag — Feature Flags · Client-Side vs Server-Side Testing.
3. Launch — ramp, then commit
For anything that could break a purchase, ramp. 10% of traffic to the variant for 48 hours, watching errors, guardrails and the split — not the primary. Then straight to 50/50.
Analyse only the full-allocation period, or weight by period. Pooling a 10/90 ramp with the 50/50 that follows mixes days where the variant was rare with days where it was common, and if conversion differs by day — it does — the pooled comparison is biased by the mix. That’s Simpson’s Paradox built into the design — Traffic Allocation. The plan says which; here, the 48h ramp is discarded and the five weeks start at 50/50.
Announce it — the name, the surface, the end date, and the one line that saves the most time: results will not be shared before the end date.
4. Run — the daily log, and what goes wrong
Do not look at the primary metric. Checking it daily and stopping when it looks good makes a 5% false-positive rate several times that — Peeking. The daily log checks everything except the answer:
DAILY LOG EXP-041
day date control variant split SRM p errors LCP p75 guardrails notes
1 10-07 2,368 2,371 50.0% 0.97 ok ok ok
2 10-08 2,410 2,395 50.2% 0.83 ok ok ok
…
9 10-15 2,288 2,306 49.8% 0.79 ok ok ok email campaign; both arms
…
The SRM check is the first column that matters. At the end of the five weeks, two possible outcomes for the same total:
control variant total expected chi-squared p
healthy 39,912 39,688 79,600 39,800 0.63 0.43
mismatch 40,212 39,388 79,600 39,800 8.53 0.003
A 50.5/49.5 split looks trivial and isn’t: on 79,600 sessions it happens by chance three times in a thousand. Something is losing variant users — a redirect, a crash, a bot filter, a cache — and the lost users aren’t random, so the comparison is broken whatever it says — Sample Ratio Mismatch.
In plain terms: a coin that lands heads 50.5% of the time over 80,000 flips isn’t a fair coin, and a split that uneven means the test isn’t measuring what you think.
Mid-test incidents, and the pre-agreed response to each:
| What happens | Response |
|---|---|
| Bug in the variant only | Fix, then restart — the data before the fix measured a different treatment. Don’t stitch the two periods together |
| Bug affecting both arms equally (site-wide outage, tracking gap) | Continue; exclude the affected window from both arms, and extend by the lost days |
| SRM detected | Stop analysis; find the cause. Almost always a restart |
| A guardrail hits its harm line once | Investigate; stop if it holds at the next check (the plan’s two-in-a-row rule) |
| An unplanned campaign or sale | Continue if it hits both arms, annotate the log, and check the result doesn’t rest on that week alone — Seasonality in Tests |
| Someone ships a change to the tested surface | Treat as a variant bug if it affects arms differently; otherwise annotate |
| A stakeholder asks how it’s going | Share the daily log — split, errors, guardrails. Not the primary. “It’s healthy; results on the 11th” |
| It’s “obviously winning” in week two | Not a reason. Stopping for success isn’t in the rules — Stopping Rules |
| A collision — another test launches on the same surface | Ask for it to wait. If it can’t, record the overlap and plan to check the combination — Interaction Effects |
The restart rule is the one people resist, because five days of data feel expensive. They’re worth nothing: they measured a treatment you’re no longer running.
5. Decide — apply the rule, don’t re-argue it
Read in the order Reading a Test Result sets out: SRM, then data quality, then guardrails, then the primary, then everything else. For ecommerce, then the four checks in Guide - Ecommerce CRO §14 — contribution, returns, customer mix, the statistical traps.
EXP-041 after five weeks at 50/50:
exposed orders conversion
control 41,320 1,653 4.000%
variant 41,330 1,769 4.280%
difference +0.280 points = +7.0% relative
standard error √(p₁(1−p₁)/n₁ + p₂(1−p₂)/n₂) = 0.139 points
95% interval 0.280 ± 1.96 × 0.139 = +0.008 to +0.551 points
= +0.2% to +13.8% relative
p 0.044
in pounds, per year 72,000 × 4.0% × lift × 12 × £23.01
point estimate +7.0% ≈ £55,600
interval +0.2% … +13.8% ≈ £1,600 … £109,600
SRM p 0.43 · guardrails clear · contribution/session +6.1% · add-to-basket −1.4% (ns)
diagnostic: delivery-step exits 9.0% → 6.2% ← the mechanism moved as predicted
What to notice. Significant, only just: the interval’s low end is almost zero. Its top end exceeds the audit’s upper bound of about 9.6%, so that part of the range isn’t plausible — the believable range is roughly +0.2% to +9.6%. And a result that only just clears a significance bar was selected partly for looking good, so expect the shipped effect to be smaller — Winner’s Curse. Take £55,600 with the usual haircut of about 30% and plan on something nearer £39,000 a year. Confidence Intervals · Practical vs Statistical Significance.
In plain terms: it probably helps; how much could be anything from almost nothing to about £75,000 a year, and the most likely value is more modest than the headline.
Applying the rules: primary up and significant, guardrails clear, diagnostic moved → SHIP. It would have been SHIP under the non-inferiority row as well, so this decision was never in doubt — what the test added is the evidence that the mechanism is real, which is the part worth keeping.
The decision record:
DECISION EXP-041 2026-11-12
Result +7.0% order conversion of exposed sessions (95% CI +0.2% to +13.8%),
p = 0.044. Guardrails clear. SRM clean.
In £ ≈ £55,600/yr point estimate; plan on ~£39,000 after winner's curse
Mechanism SUPPORTED — delivery-step exits fell 9.0% → 6.2%
Rule row 2: significant, guardrails clear, diagnostic as predicted
Decision SHIP to 100% behind the flag on 11-14
Caveats effect size imprecise; underpowered for effects below 10%
Follow-up post-launch read at 30 and 90 days — booked 12-14 and 02-12
Signed owner · analyst · approver
Write the caveats line even for a clear win. It’s the line the next person reads before building on this.
For the harder outcome: an inconclusive result, written up with the same care, is most of what a programme produces — Inconclusive Results.
6. Close — make it findable in two years
- Ship behind the flag, then remove the test code once 100% is confirmed. Dead variant code accumulates into the thing nobody dares touch
- Annotate analytics on the ship date, so next year’s trend investigation doesn’t rediscover it — Annotation and Change Logs
- Re-read at 30 and 90 days. Expect less than the test showed. Nothing at all means a false positive or an implementation that differs from the variant — Post-Test Validation
- Update the baselines that the next test’s sample size will use
- Write the archive entry:
ARCHIVE EXP-041 delivery cost on product page SHIPPED
Surface product page, all devices Dates 2026-10-07 → 11-11
Belief Surprise costs at the point of commitment cause checkout exits;
showing them earlier recovers orders rather than losing them.
Evidence: +7.0% (CI +0.2 to +13.8), diagnostic exits −2.8 points.
Confidence moderate — one test, borderline
Tags pricing-transparency · delivery · cost-surprise · PDP
Related audit F3 · EXP-039 (address lookup, same checkout step)
Next free-delivery threshold messaging; returns-cost transparency
Archive the belief, not just the outcome. “EXP-041 won” helps nobody in two years; “surprise costs at commitment cause exits, moderate confidence” is something the next person can build on, or contradict — Experiment Archive · Institutional Learning.
How tests fail operationally
- No written decision rules. The result arrives and the rule is invented to fit it
- The owner reads the result. The person who wants it to win is the worst placed to read it. Different person, or at least the pre-agreed rule
- Primary metric conditioned on a post-treatment step, so the arms compare different populations
- Peeking, and its polite form — sharing “early directional results” that then can’t be un-shared
- Stitching data across a variant fix instead of restarting
- Pooling a ramp with the full-allocation period
- No kill switch, so a harmful variant runs until the next deploy window
- Concurrent tests on one surface, discovered at the read
- Archiving the outcome without the belief, so the same idea is re-run in eighteen months
- Skipping the post-launch read, so a false positive becomes a permanent assumption
The short version
- Write the plan before the data exists: hypothesis, one primary on the exposed population, guardrails with harm lines, MDE and duration, and the decision rules
- Tick every launch-check line, including the kill switch
- Ramp for 48h if it touches the purchase; analyse only the 50/50 period
- Log daily without looking at the primary; SRM first
- Variant bug → restart. Shared bug → exclude the window from both arms. Nobody sees early results
- Apply the pre-agreed rule; state the result as a range in pounds, with the haircut
- Ship behind a flag, remove the code, re-read at 30 and 90 days, archive the belief
Related: Guide - Running a Conversion Audit for where tests come from · Guide - Statistics for CRO for the arithmetic · Guide - Ecommerce CRO for the commercial reads. Full cross-domain view: CRO.