Tags: experimentation commerce guide
Guide - Ecommerce CRO
Date: 2026-09-10
The whole job, in the order it happens: decide what you’re actually optimising, prove the numbers, model the funnel in pounds, find headroom, form a mechanism, test what can be tested, read the result including the parts the dashboard hides, and make the programme compound. Concepts link out — this guide does the sequence and the arithmetic.
Runs on one invented example throughout, so every figure downstream is checkable against the one above it: a UK direct-to-consumer retailer, 120,000 sessions a month, 2.40% session conversion rate, £62 average order value, 45% cost of goods, 12% returns.
The commercial model, built once and used everywhere below:
PER ORDER PER MONTH
goods revenue £62.00 sessions 120,000
shipping revenue £1.10 ← blended; conversion 2.40%
28% of orders 2,880
orders pay
£3.95
────────
total revenue £63.10
− cost of goods (45%) £27.90
− payment fees £ 1.40 1.9% + 20p
(1.9% × 63.10 + 0.20)
− pick and pack £ 1.20
− carriage £ 4.95
────────
contribution, order kept £27.65
contribution, order −£11.05 ← NOT zero
returned goods go back to stock;
fees, pick/pack and both
legs of carriage do not
(−1.40 −1.20 −4.95 −3.50)
blended at 12% returns:
0.88 × 27.65 + 0.12 × (−11.05) = £23.01 per order
CONTRIBUTION PER SESSION £0.552 = 23.01 × 0.024
Two things to notice. A returned order costs £11.05 rather than earning nothing — it is a negative, and at 12% it removes £1.33 from the average order on top of the £3.32 of contribution those orders never earned — £4.64 in all. And the last line, £0.552, is the number the rest of this guide optimises.
1. Decide what you’re optimising
Conversion rate is a ratio, and both halves of it move. Optimising the ratio alone is how a programme improves every reported number while the business earns less.
The identity, and where each stage of work attaches:
profit = sessions × conversion rate × contribution per order
↑ ↑ ↑
acquisition CRO's usual pricing, delivery,
owns this territory returns, product mix
— and CRO changes it
without meaning to
The third term is the one that gets ignored, and it is not fixed. A change that lowers the free-delivery threshold, adds a discount code field, promotes cheaper stock, or removes a size guide moves contribution per order while moving conversion rate the other way.
So the objective is contribution per session, not conversion rate. Contribution here is Contribution Margin — revenue minus every cost that varies with the order — applied per session so it’s comparable across tests with different traffic.
Written out once, because you will recalculate it for every result you read:
// Contribution per session — the number a CRO test should move.
// Conversion rate, AOV and revenue can all improve while this falls.
function contributionPerSession({
cr, // orders ÷ sessions
goods, shipRevenue, // £ per order
cogsRate, // cost of goods, share of goods revenue
feeRate, feeFixed, // payment processing
pickPack, carriage,
returnRate, returnCarriage,
}) {
const revenue = goods + shipRevenue;
const fees = revenue * feeRate + feeFixed;
const kept = revenue - goods * cogsRate - fees - pickPack - carriage;
// A returned order is a loss, not a zero: stock comes back,
// but everything already spent stays spent and you pay carriage twice.
const returned = -(fees + pickPack + carriage + returnCarriage);
const perOrder = kept * (1 - returnRate) + returned * returnRate;
return perOrder * cr;
}Three secondary objectives sit alongside it, and each needs naming explicitly or it gets optimised away:
| Objective | Why it can’t just be folded in | Where it’s owned |
|---|---|---|
| First-order contribution | The number above. Measurable inside a test window | This guide |
| Lifetime value | A discount that buys a worse cohort looks fine for 90 days | Customer Lifetime Value · Cohort Revenue |
| Return rate | Lands 30–60 days after the test is read | Return Rate and Reverse Logistics |
| Brand and trust | No test window contains it; only guardrails and judgement protect it | Trust Signals · Deceptive Design |
Set the primary metric before anything else. For most work this is session conversion rate as the primary metric — sensitive, fast, well understood — with contribution per session computed alongside as the decision metric. That combination is deliberate: conversion rate is what you can detect, contribution is what you act on. See Metric Design for why the denominator choice matters, and Sessionisation before trusting any session-scoped rate at all.
2. Establish scope and get access
Access chased mid-project is the largest single source of delay. Before the first analysis:
- Analytics with query access, ideally the warehouse rather than the reporting UI — Warehouse-First Analytics
- The order system, for reconciliation and for real contribution figures
- Finance’s actual numbers — cost of goods, carriage rates, payment fees, return rates by category. Estimating these is how the whole model quietly becomes wrong
- Session replay and form analytics, or the budget to add them
- A staging environment you can transact on, with test payment credentials
- The testing tool, and confirmation of what it can and cannot reach — checkout especially
- Deploy access or a named developer, with an agreed turnaround
- Someone who can authorise a price, delivery or returns-policy change. Without this, half the highest-value findings are unactionable and you should know that on day one
Then scope explicitly: which templates, which markets, which devices, and what the decision at the end actually is. “Improve conversion” is not a scope; “raise contribution per session on the mobile checkout by 5% within two quarters” is.
3. Gate one: are the numbers real
A large share of conversion problems are measurement problems, and discovering that after a quarter of testing is the classic waste. Run Guide - Auditing a Tracking Plan properly if the site is unfamiliar. The ecommerce-specific minimum, in order of how often each bites:
- Purchase reconciliation. Analytics orders and revenue against the order system, same period, same timezone. Expect a gap of a few per cent from consent denial and blockers; a gap that drifts is an open incident. Timezones and Date Boundaries first — a large share of apparent discrepancies are date boundaries
- Revenue units. Pounds or pence, one of them, everywhere. A field expecting pounds fed subunits gives revenue 100× too high, and it usually appears on one payment method only
- Duplicate purchases. Refresh the confirmation page, press back, re-enter. Anything firing twice inflates conversion rate permanently — Double Counting and Idempotency and Deduplication
- Purchase on failure. Does the order event fire when payment declines, or when an optimistic UI shows success before the server agrees? The most expensive defect class, and rarely tested
- Bots and internal traffic. Filtered, and check the filter still works. Uptime monitors, scrapers and the office all inflate the denominator — Bot and Internal Traffic
- Consent state. What proportion of sessions are unmeasured, and does that proportion differ by device or channel? If it does, every segment comparison below is contaminated — Consent Management · Ad Blockers and Tracking Loss
- Checkout coverage. On a hosted or locked-down checkout, some steps simply aren’t observable. Record it as a known limit rather than analysing around it — Checkout Instrumentation Constraints
- Identity. The stitch rate gates every new-versus-returning cut you’re about to make — Identity Stitching · Anonymous and Identified Users
- Schema. Item-level events carrying price, quantity, variant and category, consistently — Ecommerce Event Schema · Revenue Metrics
If any of this fails, fix it first. The one exception is testing: an A/B test measures the difference between two arms in the same broken instrument, so a consistent bias mostly cancels. A bias that differs between arms does not, which is why Sample Ratio Mismatch outranks everything else once tests are running.
Symptom-first entry point when a specific number looks wrong: the symptom list and Guide - Diagnosing a Metric Movement.
4. Build the funnel model in pounds
Every step, its rate, and its absolute loss. One month:
step sessions step rate lost at step
entered the site 120,000
→ viewed a product 72,000 60.0% 48,000
→ added to basket 14,400 20.0% 57,600
→ started checkout 6,912 48.0% 7,488
→ completed the order 2,880 41.7% 4,032
end to end 2.40%
Now the finding that changes how prioritisation is done, and it is the opposite of the standard advice.
In a multiplicative funnel, a 10% relative improvement is worth exactly the same at every step. The steps multiply, so improving any one of them by 10% relative improves the end-to-end rate by 10% relative, wherever it sits:
step improved by 10% relative new step rate new end rate extra orders/mo contribution/yr
viewed a product 60% → 66.0% 2.64% +288 £79,500
added to basket 20% → 22.0% 2.64% +288 £79,500
started checkout 48% → 52.8% 2.64% +288 £79,500
completed the order 41.7% → 45.8% 2.64% +288 £79,500
(288 × £23.01 × 12 = £79,513, rounded
for the small compounding differences)
The size of a drop-off tells you nothing about the value of fixing it. The 57,600 people who view a product and don’t add to basket are the biggest number on the page, and improving that step is worth precisely as much as improving the smallest one. Every “find your biggest leak” heuristic is measuring the wrong thing.
What actually differs between the steps is three things, and these are the real prioritisation inputs:
| Input | Question | Where the answer comes from |
|---|---|---|
| Headroom | How much of this step’s loss is recoverable at all? | §5 |
| Achievability | What would it cost to move it 10%? | Diagnosis and dev estimate |
| Detectability | Can you measure a 10% move here in an acceptable time? | §11 |
Detectability is where the steps genuinely diverge, and it favours the deep steps hard — see §11. Headroom favours the shallow ones, because a 60% product-view rate has more theoretical room than a 41.7% checkout completion has practical room.
Segment the funnel before leaving this step. One blended funnel hides the site. The cuts that pay in ecommerce, in order:
- Device. Mobile is usually the majority of sessions and the minority of revenue, and the gap is the single largest structural finding on most sites
- New versus returning. Two different funnels sharing a template
- Channel. Paid social converts differently from branded search because the intent differs, not because the site does — Channel Mix
- Category or product type. Considered purchases and repeat consumables behave nothing alike
- Basket size band. Where a delivery threshold or a finance option kicks in
- Landing template. Home, category, product, campaign page
See Funnel Analysis for the step-definition choices that move the answer, and Segmentation (analysis) for the group the average was hiding.
5. Find the headroom
Headroom is what a step could plausibly reach, and there are three ways to estimate it. They rank clearly.
Your own best segment. The strongest available evidence, because it holds the product, brand and price constant:
checkout start → order completion
desktop, returning 62.0%
desktop, new 48.0%
mobile, returning 44.0%
mobile, new 31.0% ← 58% of all checkout starts
─────────────────────────────────
blended 41.7%
Closing half the gap between mobile-new and desktop-new — 31% to 39.5% — on 4,009 mobile-new starts a month is +341 orders a month, ≈£94,000 a year in contribution.
4,009 × (0.395 − 0.31) = 341 orders · 341 × £23.01 × 12 = £94,157
Treat that as an upper bound, not a target. Part of the desktop-mobile gap is the population, not the experience: desktop skews toward higher-intent, higher-consideration, older, at-work-and-committed sessions. The recoverable portion is the part attributable to the interface, and nothing in the segment comparison separates the two. This is Confounding Variables in its most common commercial form, and quoting the full gap as an opportunity is the most frequent overclaim in CRO proposals.
In plain terms: mobile converting worse than desktop is partly the phone and partly the people holding it. You can fix the first and not the second, and this comparison can’t tell you the split.
Published benchmarks. Weak. Definitions differ — sessions versus users, bots filtered or not, which markets — and the composition of the sample is unknown. Useful only for a sanity check of the order of magnitude. See Benchmarking.
Theoretical ceiling. Useless as a target, occasionally useful as a bound: nobody converts 100%, and a step where most of the drop is people correctly deciding not to buy has almost no headroom regardless of what the number looks like.
Not every drop is a loss. A category page that sends 40% of visitors away may be doing its job — qualifying people out of a product that doesn’t suit them prevents a return that would have cost £11.05. Before treating a step as a leak, decide whether the people leaving would have been good orders.
6. Diagnose
Quantitative first to locate, qualitative second to explain. The other order produces confident answers to the wrong question. Qualitative vs Quantitative Research on why the two aren’t interchangeable.
The rapid pass, first hour on an unfamiliar site. Not a substitute for the work below, but it catches the embarrassing things before anyone spends a week:
- Buy something. On a phone, on mobile data, as a new customer, with a real card. Note every moment you hesitate
- Buy something as a returning customer, and again with a guest checkout
- Try to return something, or find out how
- Search for a product using the word a customer would use, not the word the site uses
- Load the product page on a throttled connection and watch what appears when
- Tab through the checkout with a keyboard only
- Read the delivery proposition. Can you tell, from the product page, when it will arrive and what it will cost?
Then the quantitative sweep, roughly in cost order:
- Funnel Analysis by segment — §4 and §5
- Form Analytics — field-level abandonment and re-entry rates. The highest-yield single diagnostic in checkout, and it names the field rather than the step
- Path Analysis — what people do instead, especially loops back to search or category
- Internal search logs — queries with no results, queries with high exit rate, and queries whose volume implies a navigation failure. See Search and Findability
- Error and validation logs — client-side validation failures by field, payment declines by method, JavaScript errors by browser. Error Tracking
- Speed by template from field data, at the 75th percentile, split by device — Core Web Vitals · Field vs Lab Data
- Stock and availability — a conversion “problem” that is actually Stockouts and Availability
Then the qualitative, aimed only at what the numbers located:
- Session Replay — watch the specific step, filtered to the specific segment. Watching randomly is entertainment
- Heatmaps — aggregate attention and interaction, with the standard caveats about scroll maps on dynamic pages
- Usability Testing — explains why faster than any amount of data. Five sessions on the located step beats a broad study
- Voice of Customer Data — support tickets, chat transcripts, product reviews and returns reasons. Returns reasons are the most under-used qualitative source in ecommerce and they arrive free
- Exit-intent and post-purchase surveys — Surveys, and the sampling caveats that come with them
- User Interviews where the question is about the decision rather than the interface
Then reconcile. The sources will disagree. Triangulation is the discipline of deciding what to believe when replay says one thing and the funnel says another — usually that both are true of different populations.
7. The surface checklist
Everything below is a consideration to check, not a prescription to apply. Each links to the note that owns it. Work the templates in the order traffic hits them.
Entry and landing
- Does the page answer why this, why you, why now, above the fold and without scrolling on a phone? Value Propositions
- Message match between the ad or email and the page. A mismatch reads as the wrong page and bounces before anything else matters — Landing Page Strategy
- Delivery, returns and price positioning visible without hunting
- Is the primary action the primary visual element? Reading Behaviour Online
- What does a first-time visitor learn about who you are? Trust Signals
- Carousels: measure whether anything past slide one is ever seen before defending it
Navigation and information architecture
- Do the top-level labels match how customers name the category, or how the business is organised internally? Taxonomy and Labelling
- Depth versus breadth, and whether a mega-menu is helping or is Choice Overload in a dropdown
- Can you tell where you are and how to get back? Wayfinding · Navigation Patterns
- Mobile navigation as a separate design problem, not a collapsed desktop one — Mobile Interaction Patterns
- Is the structure validated against how customers group things, or assumed? Card Sorting and Tree Testing · Information Architecture
Search
- Zero-result rate, and what happens on zero results — Empty States
- Synonyms, misspellings, plurals, product codes, and the words customers use that the catalogue doesn’t
- Is search prominent on mobile, or hidden behind an icon on a site where search converts at multiples of browse?
- Autocomplete: suggestions, products, or both, and whether it’s fast enough to be used
- Result relevance and default ordering — the same problem as merchandising below
Category and product listing pages
- Default sort order. It is a merchandising decision and it is rarely deliberate — Merchandising
- Filters: which facets, in which order, how many values, and whether they reflect how people choose. Faceted Filtering
- Filter behaviour on mobile — the most commonly broken interaction on a retail site
- Result count, and whether pagination, load-more or infinite scroll suits the browsing pattern
- What the card shows: price, variant availability, review count, delivery signal, second image on hover
- Out-of-stock items — surfaced, ranked down, or hidden, and what that does to the crawlable set
- Speed, because these pages are usually the heaviest — Category Page Design
- The SEO constraint on filter URLs, which is a real limit on what you’re free to change — Faceted Navigation and Crawl Budget
Product page
- Does it answer every question that would otherwise stop the purchase, in the order people ask them? Product Page Design
- Suitability and fit information — the most under-served question, and the one that drives returns
- Images: enough, zoomable, in context, showing scale, showing the variant selected
- Total cost visible: price, delivery cost, delivery date. A cost revealed later is the largest cause of abandonment, and in the UK it’s also a regulatory question
- Variant selection: what happens when a variant is out of stock, and whether the selection survives a back-navigation
- Stock and delivery urgency that is true. Fabricated scarcity is both a trust risk and an enforcement risk — Scarcity and Urgency
- Reviews: presence, volume, recency, and whether negative ones are visible. Social Proof
- Returns policy, stated on the page rather than linked into a footer
- Cross-sell and bundle placement, judged on contribution rather than on attach rate — Bundling · Basket Composition
- Copy that answers objections rather than describing features — Product Descriptions · Microcopy
Basket
- Is it a page, a drawer, or both, and does the add-to-basket confirmation interrupt browsing?
- Editable quantity, easy removal, saved for later
- Delivery cost and date shown here, not deferred to checkout
- Distance to a free-delivery threshold, if you have one, with a route to closing it — Shipping Thresholds
- Discount code field: its presence sends people out of the funnel to hunt for a code. Consider what it costs against what it earns — Discounting Strategy
- Stock revalidation, and what happens when something has sold out since it was added
- The onward action: is it obvious, and is “continue shopping” competing with it?
Checkout
- Guest checkout, genuinely available and not disguised — forced account creation is a top-three abandonment cause. Checkout Design
- Total cost complete and correct on the first screen that shows any cost
- Number of steps and fields, and whether each field earns its place. Address lookup, autofill support, sensible input types and keyboards
- Validation: inline, forgiving of formatting, and never clearing entered data — Error Prevention and Recovery · Form Design
- Payment methods the market expects, including wallets and buy-now-pay-later where the category warrants it. A missing preferred method is a silent, uniform loss
- Delivery options: named dates, not “3–5 working days”, and a visible cheapest option
- Trust and security signals at the point of card entry — Trust Signals
- Progress and system state during payment processing, where slow responses cause double submissions — Feedback and System Status
- Error recovery from a declined payment, without losing the basket
- Keyboard and screen-reader operability. This is a legal exposure as well as a conversion one — Accessible Forms · Keyboard Navigation
- What can be changed at all, given the platform. On locked-down checkouts the testable surface may be small, and finding that out before designing the test saves the sprint — Shopify
Post-purchase and account
- The confirmation page as the highest-attention moment you will ever have, usually wasted
- Order tracking that doesn’t require an account
- Returns initiation that doesn’t require an email exchange, since a hard returns process reduces the next purchase, not this one
- Account creation offered after purchase, pre-filled, rather than demanded before it
- The onward lifecycle programme, which owns the second order — Lifecycle Messaging · Repeat Purchase Rate
Cross-cutting, and check these on every template
- Speed. Treat as a conversion lever with its own budget, not an engineering concern — Performance and Conversion · Largest Contentful Paint · Interaction to Next Paint · Cumulative Layout Shift
- Third-party scripts. Usually the largest controllable cost on a retail site, and a tag manager makes them easy to add and invisible to review — Third-Party Scripts · Tag Manager Performance
- Accessibility. WCAG · Colour Contrast · Screen Readers · Accessibility Law in the UK
- Mobile-specific interaction: tap target size, sticky elements eating the viewport, keyboards covering inputs, hover-dependent behaviour
- Content and tone consistency — Voice and Tone · Content Design
- Error states and edge cases: out of stock, out of area, mid-session price change, expired basket
- Internationalisation, if you ship abroad: currency, duties, and whether the delivery promise is honest
8. The levers that aren’t interface
The largest conversion movements in ecommerce are frequently commercial rather than design, and they sit outside the usual CRO remit — which is exactly why they stay unexploited. Include them in the backlog; note which need someone else’s authority.
| Lever | Mechanism | Where it’s owned |
|---|---|---|
| Delivery speed and cost | Usually the single biggest non-price factor in the decision | Operations |
| The free-delivery threshold | Moves conversion and AOV in opposite directions — see §14 | Shipping Thresholds |
| Returns policy | Drives conversion up and margin down simultaneously | Return Rate and Reverse Logistics |
| Payment methods | A missing preferred method is a uniform, invisible loss | Finance |
| Price and offer structure | Larger effects than any interface change, and harder to test | Price Testing · Price Elasticity |
| Discount depth and cadence | A third of contribution can leave on a 10% code | Discount Impact on Margin · Promotional Cadence |
| Stock availability | No amount of optimisation converts an unavailable product | Stockouts and Availability |
| Traffic quality | A page can’t convert badly matched intent | Channel Mix |
| Site speed | Consistently under-prioritised relative to measured effect | Performance and Conversion |
Price testing deserves a specific warning. Charging different customers different prices for the same item at the same moment is commercially and reputationally risky and may be legally constrained; the usual workable forms are testing across time periods, across matched markets, or on offer structure and presentation rather than on the number itself. Price Testing covers the designs; Psychological Pricing and Price Anchoring cover the presentation levers that are safe to test.
9. Prioritise in pounds, not in points
Scoring frameworks convert guesses into numbers and the numbers inherit an authority the inputs never had — Test Prioritisation is direct about this. Their real value is forcing the conversation. Use expected value instead, because the funnel model from §4 already gives you every input.
Per candidate:
expected value = P(it works) × value if it works × (1 − decay)
cost = build + design + analysis + the traffic it occupies
Worked, for the mobile checkout finding from §5:
value if it works closing half the mobile-new gap £94,000/yr
P(it works) your own win rate, adjusted for the
strength of evidence behind it 0.35
(diagnosed from form analytics and
five usability sessions, not a hunch)
decay haircut winner's curse plus effect decay 0.30
──────────
expected value 94,000 × 0.35 × 0.70 £23,030/yr
cost 12 dev-days + design + 2.5 months of
checkout traffic ≈ £14,000
──────────
ratio 1.6×
In plain terms: you will run this test roughly three times before it works once, and when it does work it will deliver less than the test said. Both of those are priced in above, and a candidate that only clears the bar without them is not a candidate.
Two adjustments that matter more than the arithmetic:
- P(it works) is your win rate, not your confidence. Use the programme’s actual history. If you don’t have one, Win Rate and Expected Value puts mature programmes at roughly one in five to one in three, lower on already-optimised surfaces
- The cost of traffic is real and usually omitted. A test occupying the checkout for 2.5 months is 2.5 months in which no other checkout test can run cleanly. On a traffic-constrained site this is the binding constraint, and cheap-and-fast beats high-value-and-slow more often than the raw ratio suggests
Then check the list for what’s missing. Backlogs generated from diagnosis skew toward the visible. Deliberately ask whether the list contains anything on speed, on delivery proposition, on payment methods, on accessibility, and on the mobile-specific version of each item — and whether the top item is a genuinely different hypothesis from the second, or the same idea twice.
Some things should not be tested at all. Fix them and move on:
- An outright defect. If checkout throws an error on one browser, that isn’t a hypothesis
- An accessibility or legal compliance fix
- Anything where the effect is obviously large and the change is obviously cheap
- Anything you would ship regardless of the result. Running a test you’ll ignore costs traffic and buys nothing — the honest version is a non-inferiority test asking “does this hurt?“
10. Write the hypothesis
A hypothesis needs a mechanism. Without one, a result teaches you nothing that generalises, and the backlog stays a list of preferences. See Hypothesis Design.
The shape, and each part is load-bearing:
BECAUSE evidence 41% of mobile checkout starts abandon at the
delivery step, and form analytics shows the
postcode field re-entered 2.3 times on average
WE BELIEVE mechanism the field rejects valid formats, and the
error message doesn't say which format it wants
— an error-prevention failure, not a
motivation failure
SO IF WE change accept all common postcode formats, add address
lookup, and make validation forgiving
WE EXPECT prediction mobile checkout completion +8% relative
primary metric session conversion rate
guardrails order errors, delivery-address accuracy at
dispatch, page speed
decided before analysis plan and segments named in advance
Two failure modes worth naming. A hypothesis with a mechanism drawn from a behavioural principle rather than from evidence about this site is a guess with a citation — the principle explains why a mechanism could work, it isn’t evidence that this mechanism is what’s happening here. And a hypothesis whose prediction is “conversion will improve” has no failure condition, so no result can contradict it.
The mechanism usually comes from one of: an interface defect, a missing piece of information, a cost or effort the customer didn’t expect, a trust gap, a speed problem, or a decision made harder than it needs to be. The behavioural notes explain each — Cognitive Load, Choice Overload, Hick’s Law, Loss Aversion, Anchoring, Framing Effects, The Default Effect, Commitment and Consistency, Social Proof, Authority — and Deceptive Design marks where several of them stop being legitimate.
Then Pre-Registration: the primary metric, the guardrails, the planned duration, the segments you’ll look at, and the decision rule, written down before data exists.
11. Design the test
Randomisation unit: the user, effectively always. Session-level randomisation means the same person sees both versions across a multi-session purchase decision, which ecommerce purchases usually are. See Randomisation Unit and Assignment and Bucketing.
Exposure, not entry. Randomise and count only people who reach the point where the change is visible. Including everyone who entered the site in a checkout test doesn’t change the relative effect, but it destroys the power — here is the arithmetic, and it’s the most useful calculation in this section.
Detecting a 5% relative improvement, 80% power, 95% significance, on two different metrics for the same underlying change:
checkout completion site conversion
baseline p 41.67% 2.40%
target 43.75% 2.52%
absolute difference d 0.0208 0.0012
pooled p̄ 0.42708 0.0246
2 × p̄(1 − p̄) 0.48938 0.04799
n = 7.84 × 2p̄(1−p̄) ÷ d²
checkout: 7.84 × 0.48938 ÷ 0.000434 = 8,840 per arm
site: 7.84 × 0.04799 ÷ 0.00000144 = 261,278 per arm
────────────
29.6× more people needed
available per month 6,912 120,000
────────────────────────────
17.4× more available
TIME TO POWER 2.6 months 4.4 months
The 7.84 is (1.96 + 0.84)² — the significance and power constants. Full derivation and the term-by-term table in Guide - Statistics for CRO; Sample Size Calculation and Statistical Power own the concepts.
The same change is measurable in 2.6 months on one metric and 4.4 on the other, because a low base rate carries enormous relative variance. You need 29.6× as many people to see it at site level, and you only have 17.4× as many.
In plain terms: measuring a checkout change across everyone who entered the site is like listening for a whisper in a stadium — the whisper is the same volume, you’ve just added 113,000 people who aren’t part of it.
The catch, and it’s why you report both: a win on checkout completion is not automatically a win on site conversion. The change may have moved when people convert rather than whether they do. Report the deep metric as primary because it’s detectable, the site metric as secondary with its wide interval acknowledged, and require the two to at least point the same way.
Then check the calendar. Test Duration and Seasonality in Tests:
- Minimum one full week regardless of sample, ideally two — weekday and weekend buy differently
- Whole weeks only, never a partial cycle
- For considered purchases, at least one full purchase-consideration window, or you’re measuring the people who decide fast
- Avoid running across a sale, a Black Friday, a school holiday or a bank holiday weekend, and if you can’t, say so at the top of the write-up
- Anything longer than about six weeks is exposed to the site changing underneath it
When you can’t power it
Most UK ecommerce sites cannot power a classic A/B test on most of their pages, and pretending otherwise is the most common failure in the discipline. The ladder, in descending order of rigour:
- Test something bigger. Sample size scales with the square of the effect, so doubling the minimum detectable effect — the smallest change you’ve decided is worth finding — quarters the traffic needed. A 10% MDE at site level takes 1.1 months rather than 4.4. Be honest that you’re now testing a different question, and that “no detectable 10% effect” doesn’t rule out a real 4% one
- Move the metric closer to the change, as above
- Widen the exposure. Site-wide rather than one template, where the change makes sense in both
- Reduce the variance. Winsorisation and Capping on order value, and Variance Reduction where users have pre-period history
- Use a design built for continuous monitoring. Sequential Testing or Always-Valid Inference spend some power to buy the licence to stop early — mostly valuable for stopping on harm rather than on success
- Report Bayesian. Probability to Beat Control and a credible interval communicate an underpowered result more honestly than a p-value does, provided the prior is stated. It does not create information that isn’t there — Bayesian vs Frequentist
- Reframe as non-inferiority. When the real question is “can we ship this without damage”, Non-Inferiority Tests answer it with far less traffic than “is it better”
- Test demand before building. Painted Door Tests measure intent for a feature that doesn’t exist yet, at a fraction of the cost
- Ship it behind a flag with a holdout, and measure the accumulated effect of everything at once over a quarter — Holdout Groups · Feature Flags
- Decide on judgement, and say so. A documented judgement call is more honest than an underpowered test, because the test produces a number people will act on. Inconclusive Results
Test capacity is a property of your traffic, not your ambition. At 120,000 sessions a month with an average test needing ~1.5 months at full allocation, sequential running gives about eight tests a year. Getting to eighteen means running several at once on non-overlapping surfaces, which introduces Interaction Effects and needs the assignment to be independent — see Traffic Allocation. Plan the year against the arithmetic rather than against a target.
12. Build and QA
Before any traffic. Experiment QA owns the general procedure; the ecommerce-specific items:
- Both variants render, on every device class, browser and viewport in your traffic, including the browsers you don’t personally use
- The full purchase path completes in the variant, with a real transaction on staging and, if possible, one real transaction in production
- Failure paths — declined card, out of stock mid-checkout, invalid address, expired session
- Flicker. A client-side test that shows the original before swapping is measuring a worse experience than either variant, and it hurts the variant specifically — Flicker and Flash of Original Content · Client-Side vs Server-Side Testing
- Speed impact of the test itself. A testing snippet in the critical path is a confound in every test you will ever run on that site
- Assignment is recorded in analytics, so you can analyse outside the testing tool — Experiment Assignment Tracking
- Assignment is stable across sessions, devices where identified, and page loads
- Bots excluded from assignment as well as from analysis
- Interaction with other running tests, personalisation rules, and any app or plugin that modifies the same surface
- Consent interaction — does the test run for users who declined, are they assigned, and are they measured? An asymmetry here is a route to Sample Ratio Mismatch
- Cached and CDN-served pages — that the variant survives the cache layer, and that the cache key includes the variant
- A-A Tests if the machinery is new or hasn’t been checked in a year. It proves the pipeline before you trust it with a decision
13. While it runs
Don’t look at the result. A fixed-horizon test has a 5% false positive rate when checked once, at the end; every extra look is another chance to cross by luck. Peeking has the arithmetic, and the effect is large enough to change conclusions.
In plain terms: stopping the moment a test looks significant means you selected for the moment it looked good, and that moment arrives by chance in tests where nothing is happening.
Do check these daily, without looking at the primary metric:
- Sample Ratio Mismatch — a 50/50 split arriving as 52/48 on large numbers invalidates the test whatever it says. First check, every day, no exceptions
- Both variants still render. The most common failed experiment is one that quietly stopped showing the variant
- Guardrails. Order errors, payment failures, page speed, JavaScript error rate, add-to-basket rate. Guardrail Metrics
- Stock. A test whose winning variant promotes a product that sold out in week two is measuring stock, not design
Stop for harm, never for success. A pre-agreed harm threshold is the only legitimate early stop — Stopping Rules.
14. Read the result
Guide - Statistics for CRO has the ordered checks and the interval arithmetic. Run those first, then the four ecommerce-specific checks below, which are where the money is actually decided.
Check one: did contribution move
The trap, worked. A test lowers the free-delivery threshold from £60 to £40:
CONTROL VARIANT
conversion rate 2.400% 2.640% +10.0% ← primary metric
goods AOV £62.00 £58.00 −6.5%
shipping revenue, blended £1.10 £0.47 (12% of orders now pay)
──────── ────────
total revenue per order £63.10 £58.47
− cost of goods (45%) £27.90 £26.10
− payment fees £ 1.40 £ 1.31
− pick and pack £ 1.20 £ 1.20
− carriage £ 4.95 £ 4.95
──────── ────────
contribution, kept £27.65 £24.91
contribution, returned −£11.05 −£10.96
blended at 12% £23.01 £20.60
REVENUE PER SESSION £1.5144 £1.5436 +1.9% ← also up
CONTRIBUTION PER SESSION £0.5522 £0.5440 −1.5% ← the truth
≈ −£11,700 a year
Conversion rate up 10%, revenue up 1.9%, profit down. Every metric on a standard experiment dashboard reports a clear win. Only the contribution line catches it, and it catches it because the extra orders arrive smaller while carriage stays fixed at £4.95 whatever the basket is worth.
This is not an exotic case. It is the default outcome of any change that trades basket size for order count — thresholds, discount codes, entry-price product promotion, payment plans, and most bundle placements. See Shipping Thresholds and Discount Impact on Margin.
Check two: what happened to returns
Returns land 30–60 days after the order, which is usually after the test has been read and shipped. A change that removes friction from a decision people should have taken more slowly raises the return rate, and the £11.05 negative from the opening model does the rest:
same site, a variant that raises conversion 3% and returns 12% → 15%
contribution per order 0.85 × 27.65 + 0.15 × (−11.05) = £21.84 (was £23.01)
contribution per session 21.84 × 0.02472 = £0.5400 (was £0.5522)
−2.2%
≈ −£17,500 a year
So return rate is a guardrail that cannot be read inside the test window. Two workable responses: keep the assignment recorded and re-read the cohort at day 60, or refuse to ship anything that touches sizing, suitability, product information or returns policy without that lagged read. The first is better and requires only that assignment lives in the warehouse.
Check three: which customers, and what happens next
- New versus returning. A win driven entirely by returning customers may be a change in timing rather than in demand
- Category mix. Conversion up because cheaper stock got promoted is a merchandising change wearing a test’s clothes — Merchandising · Basket Composition
- Cohort quality. Discount-acquired customers repeat at lower rates. A first-order win can be an LTV loss, and the test window can’t see it — Cohort Analysis · Customer Lifetime Value
- Cannibalisation. Did the winning surface take orders from another surface rather than create them? Site-level conversion is the check
Check four: the standard statistical traps
Ranked by how often they actually bite in commercial testing:
- You looked early and stopped — Peeking
- You reported the metric that won out of several tracked — The Multiple Comparisons Problem
- You found the win in a segment after the fact. Post-hoc slicing produces significance from noise reliably; if a segment mattered it was named before launch — Segmentation (test results)
- The shipped effect underdelivers. Expected, by construction: you selected the winner on a noisy estimate, so the estimate was biased high the moment you chose it — Winner’s Curse
- The effect faded. Novelty and Primacy Effects if regulars reacted to the change itself, Regression to the Mean if you intervened on something already at an extreme
- Every segment improved and the total didn’t — Simpson’s Paradox, traffic mix moved
- It’s significant but not worth building — Practical vs Statistical Significance
Then state the result as a range in pounds, with a recommendation. Communicating Uncertainty and Reading a Test Result.
15. Ship, and check it landed
- Ship behind a flag, so the winner goes live without a release and comes back off without one — Feature Flags · Progressive Delivery
- Remove the test code. Accumulated dead variant code is a real and under-discussed cost on long-running programmes
- Re-measure at 30 and 90 days. The shipped effect should be smaller than the test effect; if it’s absent, the test was a false positive or the implementation differs from the variant — Post-Test Validation
- Re-read returns and cohort quality at 60 days, per §14
- Archive the result, including losses and inconclusives, in a form someone can search in two years. Experiment Archive · Institutional Learning
- Annotate the analytics, so next year’s trend investigation doesn’t rediscover this as an anomaly — Annotation and Change Logs
- Update the funnel model. The baselines in §4 have moved, and every future prioritisation uses them
16. Make the programme compound
Individual tests barely matter. The system that produces them does.
The programme’s actual return, run for our example site at £795,000 annual contribution:
tests run in a year 18
win rate 22% → 4 winners
average validated lift of a winner +3% → £23,900/yr each
gross value of winners 4 × 23,900 = £95,600
winner's curse and decay haircut × 0.70 = £66,900
losses prevented 14% of 18 = 2.5 tests at −2.4%
2.5 × 19,100 = £47,750
← only counts if you would
genuinely have shipped them
programme cost tooling, analysis time,
design and dev ≈ £55,000
─────────
net, winners alone +£11,900
net, including prevented losses +£59,650
The second line is the one that decides whether the programme is worth running, and it is the line nobody reports. Prevented losses are invisible by construction: nothing happened. Win Rate and Expected Value argues this at length, and it is the argument that keeps programmes funded through a bad quarter.
Verify the whole thing with a holdout. A 5% holdout excluded from every shipped change for a year is the only measurement of the programme’s real cumulative effect, and it routinely shows less than the sum of the individual test results. That gap is the honest number.
Then the things that raise the return more than any individual test:
- Experimentation Velocity — tests per period is the dominant term in the arithmetic above, and it’s limited by traffic, dev capacity and decision latency rather than by ideas
- Institutional Learning — turning results into beliefs about customers, and revising the beliefs when results contradict them. Without this you retest the same idea every eighteen months
- Experimentation Maturity — identifying which stage is actually the bottleneck before investing in the others
- Experiment Archive — searchable, including failures, with the hypothesis and the mechanism preserved rather than just the outcome
17. Conduct, law and the line
Not optional, and cheaper to consider now than after an enforcement letter.
- Total price disclosure. Deferring a mandatory cost to a later step is both the largest abandonment cause and a regulatory question in the UK. Deceptive Design · Checkout Design
- Urgency and scarcity claims must be true. A countdown that resets and a “3 left” that never changes are enforcement risks, not persuasion techniques — Scarcity and Urgency
- Reviews. Fabrication and suppression are both live enforcement areas — Social Proof
- Defaults and pre-ticked boxes, especially where they opt someone into a subscription or a cost — The Default Effect · Commitment and Consistency
- Consent, and whether your test tool runs before it — Consent Management · UK GDPR and PECR for Analytics
- Accessibility, which is a legal exposure as well as a conversion one — Accessibility Law in the UK
- Experimenting on people — what’s reasonable to test on customers who didn’t consent to being in a study, and where the line sits for price, for vulnerability and for harm. Ethics of Experimentation · Testing and Compliance
[CHECK: current CMA position on drip pricing, urgency claims and review practices, and the operative consumer-protection regime, before relying on any specific statement of what is prohibited. This area has been actively legislated and enforced, and the detail matters more than any summary.]
18. The first ninety days
A defensible sequence when you take over an unfamiliar site.
Weeks 1–2 — trust the data. Access. Reconcile orders and revenue against finance. Purchase-event QA. Build the opening commercial model with finance’s real numbers. Output: a measurement findings list, and a model you can price things with.
Weeks 3–5 — see the shape. Funnel by segment. Headroom analysis. The rapid pass and the quantitative sweep. Speed audit. Three to five usability sessions on whatever the numbers located. Output: a prioritised backlog in pounds, and a written view of where the money is.
Weeks 6–8 — prove the machinery. An A/A test. Fix the outright defects found in weeks 3–5 without testing them. Ship the accessibility and legal items. Launch the first real test. Output: a working pipeline and something in flight.
Weeks 9–13 — establish the rhythm. Second and third tests. Set up the holdout. Write the archive template and the result write-up format. First quarterly report, including the losses prevented. Output: a programme rather than a project.
Ship the defects immediately and don’t wait for the testing programme. A quarter of measurement work with nothing shipped is how CRO engagements lose their sponsor in month two.
19. How programmes fail
Ordered by frequency, and each has a section above that prevents it.
- Optimising conversion rate while contribution falls. §1 and §14
- Testing on traffic that can’t detect anything, then acting on the numbers anyway. §11
- Building the backlog from opinion rather than diagnosis. §6
- Testing tiny changes because they’re easy to ship, which produces a year of underpowered nulls. §9
- Never touching the commercial levers, because they belong to someone else. §8
- Reading results early, and reading them charitably. §13
- Not counting the losses prevented, so the programme looks like it isn’t working. §16
- No archive, so the same test recurs every eighteen months. §15
- Measurement never gated, so a quarter of results are artefacts. §3
- Optimising the first order and ignoring the second. §1 and §14
The short version
- The objective is contribution per session. Conversion rate, revenue and AOV can all rise while it falls, and that is the normal outcome of threshold, discount and bundling changes
- Prove the numbers before analysing them, and expect a measurable share of “conversion problems” to be tracking problems
- Drop-off size doesn’t indicate value. In a multiplicative funnel a 10% improvement is worth the same at every step; what differs is headroom, achievability and detectability
- Calculate the sample size before designing anything. Most tests are unaffordable and it takes ten minutes to find out
- Measure as close to the change as the question allows. The same effect can be 1.7× faster to detect on a deeper metric
- A returned order costs money, and the return arrives after the test is read
- Don’t test what you’d ship anyway, and don’t test outright defects. Fix them
- Count the losses you prevented. It’s half the programme’s value and none of its reporting
Related, in the order you’d reach for them: Guide - Statistics for CRO for the arithmetic of a single test · Guide - Running an Experiment for the procedure around it · Guide - Running a Conversion Audit for the diagnostic sweep in §6 at full depth · Guide - Unit Economics and Commercial Decisions for the model at the top · Guide - Diagnosing and Fixing Performance for §8’s speed lever. Full cross-domain view: CRO.