Tags: experimentation commerce guide

Guide - Ecommerce CRO

Date: 2026-09-10


The whole job, in the order it happens: decide what you’re actually optimising, prove the numbers, model the funnel in pounds, find headroom, form a mechanism, test what can be tested, read the result including the parts the dashboard hides, and make the programme compound. Concepts link out — this guide does the sequence and the arithmetic.


Runs on one invented example throughout, so every figure downstream is checkable against the one above it: a UK direct-to-consumer retailer, 120,000 sessions a month, 2.40% session conversion rate, £62 average order value, 45% cost of goods, 12% returns.

The commercial model, built once and used everywhere below:

PER ORDER                                        PER MONTH

goods revenue              £62.00                sessions        120,000
shipping revenue            £1.10  ← blended;    conversion         2.40%
                                     28% of      orders             2,880
                                     orders pay
                                     £3.95
                          ────────
total revenue              £63.10

− cost of goods (45%)      £27.90
− payment fees             £ 1.40   1.9% + 20p
  (1.9% × 63.10 + 0.20)
− pick and pack            £ 1.20
− carriage                 £ 4.95
                          ────────
contribution, order kept   £27.65

contribution, order        −£11.05  ← NOT zero
  returned                            goods go back to stock;
                                      fees, pick/pack and both
                                      legs of carriage do not
  (−1.40 −1.20 −4.95 −3.50)

blended at 12% returns:
  0.88 × 27.65  +  0.12 × (−11.05)  =  £23.01 per order

CONTRIBUTION PER SESSION   £0.552    = 23.01 × 0.024

Two things to notice. A returned order costs £11.05 rather than earning nothing — it is a negative, and at 12% it removes £1.33 from the average order on top of the £3.32 of contribution those orders never earned — £4.64 in all. And the last line, £0.552, is the number the rest of this guide optimises.

1. Decide what you’re optimising

Conversion rate is a ratio, and both halves of it move. Optimising the ratio alone is how a programme improves every reported number while the business earns less.

The identity, and where each stage of work attaches:

profit  =  sessions  ×  conversion rate  ×  contribution per order

           ↑             ↑                   ↑
           acquisition   CRO's usual          pricing, delivery,
           owns this     territory            returns, product mix
                                              — and CRO changes it
                                                without meaning to

The third term is the one that gets ignored, and it is not fixed. A change that lowers the free-delivery threshold, adds a discount code field, promotes cheaper stock, or removes a size guide moves contribution per order while moving conversion rate the other way.

So the objective is contribution per session, not conversion rate. Contribution here is Contribution Margin — revenue minus every cost that varies with the order — applied per session so it’s comparable across tests with different traffic.

Written out once, because you will recalculate it for every result you read:

// Contribution per session — the number a CRO test should move.
// Conversion rate, AOV and revenue can all improve while this falls.
function contributionPerSession({
  cr,                          // orders ÷ sessions
  goods, shipRevenue,          // £ per order
  cogsRate,                    // cost of goods, share of goods revenue
  feeRate, feeFixed,           // payment processing
  pickPack, carriage,
  returnRate, returnCarriage,
}) {
  const revenue = goods + shipRevenue;
  const fees    = revenue * feeRate + feeFixed;
  const kept    = revenue - goods * cogsRate - fees - pickPack - carriage;
 
  // A returned order is a loss, not a zero: stock comes back,
  // but everything already spent stays spent and you pay carriage twice.
  const returned = -(fees + pickPack + carriage + returnCarriage);
 
  const perOrder = kept * (1 - returnRate) + returned * returnRate;
  return perOrder * cr;
}

Three secondary objectives sit alongside it, and each needs naming explicitly or it gets optimised away:

ObjectiveWhy it can’t just be folded inWhere it’s owned
First-order contributionThe number above. Measurable inside a test windowThis guide
Lifetime valueA discount that buys a worse cohort looks fine for 90 daysCustomer Lifetime Value · Cohort Revenue
Return rateLands 30–60 days after the test is readReturn Rate and Reverse Logistics
Brand and trustNo test window contains it; only guardrails and judgement protect itTrust Signals · Deceptive Design

Set the primary metric before anything else. For most work this is session conversion rate as the primary metric — sensitive, fast, well understood — with contribution per session computed alongside as the decision metric. That combination is deliberate: conversion rate is what you can detect, contribution is what you act on. See Metric Design for why the denominator choice matters, and Sessionisation before trusting any session-scoped rate at all.

2. Establish scope and get access

Access chased mid-project is the largest single source of delay. Before the first analysis:

  • Analytics with query access, ideally the warehouse rather than the reporting UI — Warehouse-First Analytics
  • The order system, for reconciliation and for real contribution figures
  • Finance’s actual numbers — cost of goods, carriage rates, payment fees, return rates by category. Estimating these is how the whole model quietly becomes wrong
  • Session replay and form analytics, or the budget to add them
  • A staging environment you can transact on, with test payment credentials
  • The testing tool, and confirmation of what it can and cannot reach — checkout especially
  • Deploy access or a named developer, with an agreed turnaround
  • Someone who can authorise a price, delivery or returns-policy change. Without this, half the highest-value findings are unactionable and you should know that on day one

Then scope explicitly: which templates, which markets, which devices, and what the decision at the end actually is. “Improve conversion” is not a scope; “raise contribution per session on the mobile checkout by 5% within two quarters” is.

3. Gate one: are the numbers real

A large share of conversion problems are measurement problems, and discovering that after a quarter of testing is the classic waste. Run Guide - Auditing a Tracking Plan properly if the site is unfamiliar. The ecommerce-specific minimum, in order of how often each bites:

  • Purchase reconciliation. Analytics orders and revenue against the order system, same period, same timezone. Expect a gap of a few per cent from consent denial and blockers; a gap that drifts is an open incident. Timezones and Date Boundaries first — a large share of apparent discrepancies are date boundaries
  • Revenue units. Pounds or pence, one of them, everywhere. A field expecting pounds fed subunits gives revenue 100× too high, and it usually appears on one payment method only
  • Duplicate purchases. Refresh the confirmation page, press back, re-enter. Anything firing twice inflates conversion rate permanently — Double Counting and Idempotency and Deduplication
  • Purchase on failure. Does the order event fire when payment declines, or when an optimistic UI shows success before the server agrees? The most expensive defect class, and rarely tested
  • Bots and internal traffic. Filtered, and check the filter still works. Uptime monitors, scrapers and the office all inflate the denominator — Bot and Internal Traffic
  • Consent state. What proportion of sessions are unmeasured, and does that proportion differ by device or channel? If it does, every segment comparison below is contaminated — Consent Management · Ad Blockers and Tracking Loss
  • Checkout coverage. On a hosted or locked-down checkout, some steps simply aren’t observable. Record it as a known limit rather than analysing around it — Checkout Instrumentation Constraints
  • Identity. The stitch rate gates every new-versus-returning cut you’re about to make — Identity Stitching · Anonymous and Identified Users
  • Schema. Item-level events carrying price, quantity, variant and category, consistently — Ecommerce Event Schema · Revenue Metrics

If any of this fails, fix it first. The one exception is testing: an A/B test measures the difference between two arms in the same broken instrument, so a consistent bias mostly cancels. A bias that differs between arms does not, which is why Sample Ratio Mismatch outranks everything else once tests are running.

Symptom-first entry point when a specific number looks wrong: the symptom list and Guide - Diagnosing a Metric Movement.

4. Build the funnel model in pounds

Every step, its rate, and its absolute loss. One month:

step                        sessions   step rate   lost at step

entered the site             120,000
  → viewed a product          72,000      60.0%      48,000
  → added to basket           14,400      20.0%      57,600
  → started checkout           6,912      48.0%       7,488
  → completed the order        2,880      41.7%       4,032

end to end                                 2.40%

Now the finding that changes how prioritisation is done, and it is the opposite of the standard advice.

In a multiplicative funnel, a 10% relative improvement is worth exactly the same at every step. The steps multiply, so improving any one of them by 10% relative improves the end-to-end rate by 10% relative, wherever it sits:

step improved by 10% relative   new step rate   new end rate   extra orders/mo   contribution/yr

viewed a product                60% → 66.0%         2.64%            +288           £79,500
added to basket                 20% → 22.0%         2.64%            +288           £79,500
started checkout                48% → 52.8%         2.64%            +288           £79,500
completed the order            41.7% → 45.8%        2.64%            +288           £79,500

                                                    (288 × £23.01 × 12 = £79,513, rounded
                                                     for the small compounding differences)

The size of a drop-off tells you nothing about the value of fixing it. The 57,600 people who view a product and don’t add to basket are the biggest number on the page, and improving that step is worth precisely as much as improving the smallest one. Every “find your biggest leak” heuristic is measuring the wrong thing.

What actually differs between the steps is three things, and these are the real prioritisation inputs:

InputQuestionWhere the answer comes from
HeadroomHow much of this step’s loss is recoverable at all?§5
AchievabilityWhat would it cost to move it 10%?Diagnosis and dev estimate
DetectabilityCan you measure a 10% move here in an acceptable time?§11

Detectability is where the steps genuinely diverge, and it favours the deep steps hard — see §11. Headroom favours the shallow ones, because a 60% product-view rate has more theoretical room than a 41.7% checkout completion has practical room.

Segment the funnel before leaving this step. One blended funnel hides the site. The cuts that pay in ecommerce, in order:

  • Device. Mobile is usually the majority of sessions and the minority of revenue, and the gap is the single largest structural finding on most sites
  • New versus returning. Two different funnels sharing a template
  • Channel. Paid social converts differently from branded search because the intent differs, not because the site does — Channel Mix
  • Category or product type. Considered purchases and repeat consumables behave nothing alike
  • Basket size band. Where a delivery threshold or a finance option kicks in
  • Landing template. Home, category, product, campaign page

See Funnel Analysis for the step-definition choices that move the answer, and Segmentation (analysis) for the group the average was hiding.

5. Find the headroom

Headroom is what a step could plausibly reach, and there are three ways to estimate it. They rank clearly.

Your own best segment. The strongest available evidence, because it holds the product, brand and price constant:

checkout start → order completion

  desktop, returning          62.0%
  desktop, new                48.0%
  mobile, returning           44.0%
  mobile, new                 31.0%   ← 58% of all checkout starts
  ─────────────────────────────────
  blended                     41.7%

Closing half the gap between mobile-new and desktop-new — 31% to 39.5% — on 4,009 mobile-new starts a month is +341 orders a month, ≈£94,000 a year in contribution.

4,009 × (0.395 − 0.31) = 341 orders · 341 × £23.01 × 12 = £94,157

Treat that as an upper bound, not a target. Part of the desktop-mobile gap is the population, not the experience: desktop skews toward higher-intent, higher-consideration, older, at-work-and-committed sessions. The recoverable portion is the part attributable to the interface, and nothing in the segment comparison separates the two. This is Confounding Variables in its most common commercial form, and quoting the full gap as an opportunity is the most frequent overclaim in CRO proposals.

In plain terms: mobile converting worse than desktop is partly the phone and partly the people holding it. You can fix the first and not the second, and this comparison can’t tell you the split.

Published benchmarks. Weak. Definitions differ — sessions versus users, bots filtered or not, which markets — and the composition of the sample is unknown. Useful only for a sanity check of the order of magnitude. See Benchmarking.

Theoretical ceiling. Useless as a target, occasionally useful as a bound: nobody converts 100%, and a step where most of the drop is people correctly deciding not to buy has almost no headroom regardless of what the number looks like.

Not every drop is a loss. A category page that sends 40% of visitors away may be doing its job — qualifying people out of a product that doesn’t suit them prevents a return that would have cost £11.05. Before treating a step as a leak, decide whether the people leaving would have been good orders.

6. Diagnose

Quantitative first to locate, qualitative second to explain. The other order produces confident answers to the wrong question. Qualitative vs Quantitative Research on why the two aren’t interchangeable.

The rapid pass, first hour on an unfamiliar site. Not a substitute for the work below, but it catches the embarrassing things before anyone spends a week:

  • Buy something. On a phone, on mobile data, as a new customer, with a real card. Note every moment you hesitate
  • Buy something as a returning customer, and again with a guest checkout
  • Try to return something, or find out how
  • Search for a product using the word a customer would use, not the word the site uses
  • Load the product page on a throttled connection and watch what appears when
  • Tab through the checkout with a keyboard only
  • Read the delivery proposition. Can you tell, from the product page, when it will arrive and what it will cost?

Then the quantitative sweep, roughly in cost order:

  • Funnel Analysis by segment — §4 and §5
  • Form Analytics — field-level abandonment and re-entry rates. The highest-yield single diagnostic in checkout, and it names the field rather than the step
  • Path Analysis — what people do instead, especially loops back to search or category
  • Internal search logs — queries with no results, queries with high exit rate, and queries whose volume implies a navigation failure. See Search and Findability
  • Error and validation logs — client-side validation failures by field, payment declines by method, JavaScript errors by browser. Error Tracking
  • Speed by template from field data, at the 75th percentile, split by device — Core Web Vitals · Field vs Lab Data
  • Stock and availability — a conversion “problem” that is actually Stockouts and Availability

Then the qualitative, aimed only at what the numbers located:

  • Session Replay — watch the specific step, filtered to the specific segment. Watching randomly is entertainment
  • Heatmaps — aggregate attention and interaction, with the standard caveats about scroll maps on dynamic pages
  • Usability Testing — explains why faster than any amount of data. Five sessions on the located step beats a broad study
  • Voice of Customer Data — support tickets, chat transcripts, product reviews and returns reasons. Returns reasons are the most under-used qualitative source in ecommerce and they arrive free
  • Exit-intent and post-purchase surveys — Surveys, and the sampling caveats that come with them
  • User Interviews where the question is about the decision rather than the interface

Then reconcile. The sources will disagree. Triangulation is the discipline of deciding what to believe when replay says one thing and the funnel says another — usually that both are true of different populations.

7. The surface checklist

Everything below is a consideration to check, not a prescription to apply. Each links to the note that owns it. Work the templates in the order traffic hits them.

Entry and landing

  • Does the page answer why this, why you, why now, above the fold and without scrolling on a phone? Value Propositions
  • Message match between the ad or email and the page. A mismatch reads as the wrong page and bounces before anything else matters — Landing Page Strategy
  • Delivery, returns and price positioning visible without hunting
  • Is the primary action the primary visual element? Reading Behaviour Online
  • What does a first-time visitor learn about who you are? Trust Signals
  • Carousels: measure whether anything past slide one is ever seen before defending it

Navigation and information architecture

Search

  • Zero-result rate, and what happens on zero results — Empty States
  • Synonyms, misspellings, plurals, product codes, and the words customers use that the catalogue doesn’t
  • Is search prominent on mobile, or hidden behind an icon on a site where search converts at multiples of browse?
  • Autocomplete: suggestions, products, or both, and whether it’s fast enough to be used
  • Result relevance and default ordering — the same problem as merchandising below

Category and product listing pages

  • Default sort order. It is a merchandising decision and it is rarely deliberate — Merchandising
  • Filters: which facets, in which order, how many values, and whether they reflect how people choose. Faceted Filtering
  • Filter behaviour on mobile — the most commonly broken interaction on a retail site
  • Result count, and whether pagination, load-more or infinite scroll suits the browsing pattern
  • What the card shows: price, variant availability, review count, delivery signal, second image on hover
  • Out-of-stock items — surfaced, ranked down, or hidden, and what that does to the crawlable set
  • Speed, because these pages are usually the heaviest — Category Page Design
  • The SEO constraint on filter URLs, which is a real limit on what you’re free to change — Faceted Navigation and Crawl Budget

Product page

  • Does it answer every question that would otherwise stop the purchase, in the order people ask them? Product Page Design
  • Suitability and fit information — the most under-served question, and the one that drives returns
  • Images: enough, zoomable, in context, showing scale, showing the variant selected
  • Total cost visible: price, delivery cost, delivery date. A cost revealed later is the largest cause of abandonment, and in the UK it’s also a regulatory question
  • Variant selection: what happens when a variant is out of stock, and whether the selection survives a back-navigation
  • Stock and delivery urgency that is true. Fabricated scarcity is both a trust risk and an enforcement risk — Scarcity and Urgency
  • Reviews: presence, volume, recency, and whether negative ones are visible. Social Proof
  • Returns policy, stated on the page rather than linked into a footer
  • Cross-sell and bundle placement, judged on contribution rather than on attach rate — Bundling · Basket Composition
  • Copy that answers objections rather than describing features — Product Descriptions · Microcopy

Basket

  • Is it a page, a drawer, or both, and does the add-to-basket confirmation interrupt browsing?
  • Editable quantity, easy removal, saved for later
  • Delivery cost and date shown here, not deferred to checkout
  • Distance to a free-delivery threshold, if you have one, with a route to closing it — Shipping Thresholds
  • Discount code field: its presence sends people out of the funnel to hunt for a code. Consider what it costs against what it earns — Discounting Strategy
  • Stock revalidation, and what happens when something has sold out since it was added
  • The onward action: is it obvious, and is “continue shopping” competing with it?

Checkout

  • Guest checkout, genuinely available and not disguised — forced account creation is a top-three abandonment cause. Checkout Design
  • Total cost complete and correct on the first screen that shows any cost
  • Number of steps and fields, and whether each field earns its place. Address lookup, autofill support, sensible input types and keyboards
  • Validation: inline, forgiving of formatting, and never clearing entered data — Error Prevention and Recovery · Form Design
  • Payment methods the market expects, including wallets and buy-now-pay-later where the category warrants it. A missing preferred method is a silent, uniform loss
  • Delivery options: named dates, not “3–5 working days”, and a visible cheapest option
  • Trust and security signals at the point of card entry — Trust Signals
  • Progress and system state during payment processing, where slow responses cause double submissions — Feedback and System Status
  • Error recovery from a declined payment, without losing the basket
  • Keyboard and screen-reader operability. This is a legal exposure as well as a conversion one — Accessible Forms · Keyboard Navigation
  • What can be changed at all, given the platform. On locked-down checkouts the testable surface may be small, and finding that out before designing the test saves the sprint — Shopify

Post-purchase and account

  • The confirmation page as the highest-attention moment you will ever have, usually wasted
  • Order tracking that doesn’t require an account
  • Returns initiation that doesn’t require an email exchange, since a hard returns process reduces the next purchase, not this one
  • Account creation offered after purchase, pre-filled, rather than demanded before it
  • The onward lifecycle programme, which owns the second order — Lifecycle Messaging · Repeat Purchase Rate

Cross-cutting, and check these on every template

8. The levers that aren’t interface

The largest conversion movements in ecommerce are frequently commercial rather than design, and they sit outside the usual CRO remit — which is exactly why they stay unexploited. Include them in the backlog; note which need someone else’s authority.

LeverMechanismWhere it’s owned
Delivery speed and costUsually the single biggest non-price factor in the decisionOperations
The free-delivery thresholdMoves conversion and AOV in opposite directions — see §14Shipping Thresholds
Returns policyDrives conversion up and margin down simultaneouslyReturn Rate and Reverse Logistics
Payment methodsA missing preferred method is a uniform, invisible lossFinance
Price and offer structureLarger effects than any interface change, and harder to testPrice Testing · Price Elasticity
Discount depth and cadenceA third of contribution can leave on a 10% codeDiscount Impact on Margin · Promotional Cadence
Stock availabilityNo amount of optimisation converts an unavailable productStockouts and Availability
Traffic qualityA page can’t convert badly matched intentChannel Mix
Site speedConsistently under-prioritised relative to measured effectPerformance and Conversion

Price testing deserves a specific warning. Charging different customers different prices for the same item at the same moment is commercially and reputationally risky and may be legally constrained; the usual workable forms are testing across time periods, across matched markets, or on offer structure and presentation rather than on the number itself. Price Testing covers the designs; Psychological Pricing and Price Anchoring cover the presentation levers that are safe to test.

9. Prioritise in pounds, not in points

Scoring frameworks convert guesses into numbers and the numbers inherit an authority the inputs never had — Test Prioritisation is direct about this. Their real value is forcing the conversation. Use expected value instead, because the funnel model from §4 already gives you every input.

Per candidate:

expected value  =  P(it works)  ×  value if it works  ×  (1 − decay)
cost            =  build + design + analysis + the traffic it occupies

Worked, for the mobile checkout finding from §5:

value if it works      closing half the mobile-new gap        £94,000/yr
P(it works)            your own win rate, adjusted for the
                       strength of evidence behind it              0.35
                       (diagnosed from form analytics and
                        five usability sessions, not a hunch)
decay haircut          winner's curse plus effect decay            0.30
                                                              ──────────
expected value         94,000 × 0.35 × 0.70                    £23,030/yr

cost                   12 dev-days + design + 2.5 months of
                       checkout traffic                        ≈ £14,000
                                                              ──────────
ratio                                                               1.6×

In plain terms: you will run this test roughly three times before it works once, and when it does work it will deliver less than the test said. Both of those are priced in above, and a candidate that only clears the bar without them is not a candidate.

Two adjustments that matter more than the arithmetic:

  • P(it works) is your win rate, not your confidence. Use the programme’s actual history. If you don’t have one, Win Rate and Expected Value puts mature programmes at roughly one in five to one in three, lower on already-optimised surfaces
  • The cost of traffic is real and usually omitted. A test occupying the checkout for 2.5 months is 2.5 months in which no other checkout test can run cleanly. On a traffic-constrained site this is the binding constraint, and cheap-and-fast beats high-value-and-slow more often than the raw ratio suggests

Then check the list for what’s missing. Backlogs generated from diagnosis skew toward the visible. Deliberately ask whether the list contains anything on speed, on delivery proposition, on payment methods, on accessibility, and on the mobile-specific version of each item — and whether the top item is a genuinely different hypothesis from the second, or the same idea twice.

Some things should not be tested at all. Fix them and move on:

  • An outright defect. If checkout throws an error on one browser, that isn’t a hypothesis
  • An accessibility or legal compliance fix
  • Anything where the effect is obviously large and the change is obviously cheap
  • Anything you would ship regardless of the result. Running a test you’ll ignore costs traffic and buys nothing — the honest version is a non-inferiority test asking “does this hurt?“

10. Write the hypothesis

A hypothesis needs a mechanism. Without one, a result teaches you nothing that generalises, and the backlog stays a list of preferences. See Hypothesis Design.

The shape, and each part is load-bearing:

BECAUSE     evidence            41% of mobile checkout starts abandon at the
                                delivery step, and form analytics shows the
                                postcode field re-entered 2.3 times on average

WE BELIEVE  mechanism           the field rejects valid formats, and the
                                error message doesn't say which format it wants
                                — an error-prevention failure, not a
                                motivation failure

SO IF WE    change              accept all common postcode formats, add address
                                lookup, and make validation forgiving

WE EXPECT   prediction          mobile checkout completion +8% relative
            primary metric      session conversion rate
            guardrails          order errors, delivery-address accuracy at
                                dispatch, page speed
            decided before      analysis plan and segments named in advance

Two failure modes worth naming. A hypothesis with a mechanism drawn from a behavioural principle rather than from evidence about this site is a guess with a citation — the principle explains why a mechanism could work, it isn’t evidence that this mechanism is what’s happening here. And a hypothesis whose prediction is “conversion will improve” has no failure condition, so no result can contradict it.

The mechanism usually comes from one of: an interface defect, a missing piece of information, a cost or effort the customer didn’t expect, a trust gap, a speed problem, or a decision made harder than it needs to be. The behavioural notes explain each — Cognitive Load, Choice Overload, Hick’s Law, Loss Aversion, Anchoring, Framing Effects, The Default Effect, Commitment and Consistency, Social Proof, Authority — and Deceptive Design marks where several of them stop being legitimate.

Then Pre-Registration: the primary metric, the guardrails, the planned duration, the segments you’ll look at, and the decision rule, written down before data exists.

11. Design the test

Randomisation unit: the user, effectively always. Session-level randomisation means the same person sees both versions across a multi-session purchase decision, which ecommerce purchases usually are. See Randomisation Unit and Assignment and Bucketing.

Exposure, not entry. Randomise and count only people who reach the point where the change is visible. Including everyone who entered the site in a checkout test doesn’t change the relative effect, but it destroys the power — here is the arithmetic, and it’s the most useful calculation in this section.

Detecting a 5% relative improvement, 80% power, 95% significance, on two different metrics for the same underlying change:

                                   checkout completion      site conversion

baseline p                                    41.67%                  2.40%
target                                        43.75%                  2.52%
absolute difference d                        0.0208                 0.0012
pooled p̄                                    0.42708                 0.0246
2 × p̄(1 − p̄)                              0.48938                0.04799

n = 7.84 × 2p̄(1−p̄) ÷ d²
    checkout:  7.84 × 0.48938 ÷ 0.000434    =  8,840 per arm
    site:      7.84 × 0.04799 ÷ 0.00000144  = 261,278 per arm
                                               ────────────
                                               29.6× more people needed

available per month                            6,912               120,000
                                               ────────────────────────────
                                               17.4× more available

TIME TO POWER                             2.6 months            4.4 months

The 7.84 is (1.96 + 0.84)² — the significance and power constants. Full derivation and the term-by-term table in Guide - Statistics for CRO; Sample Size Calculation and Statistical Power own the concepts.

The same change is measurable in 2.6 months on one metric and 4.4 on the other, because a low base rate carries enormous relative variance. You need 29.6× as many people to see it at site level, and you only have 17.4× as many.

In plain terms: measuring a checkout change across everyone who entered the site is like listening for a whisper in a stadium — the whisper is the same volume, you’ve just added 113,000 people who aren’t part of it.

The catch, and it’s why you report both: a win on checkout completion is not automatically a win on site conversion. The change may have moved when people convert rather than whether they do. Report the deep metric as primary because it’s detectable, the site metric as secondary with its wide interval acknowledged, and require the two to at least point the same way.

Then check the calendar. Test Duration and Seasonality in Tests:

  • Minimum one full week regardless of sample, ideally two — weekday and weekend buy differently
  • Whole weeks only, never a partial cycle
  • For considered purchases, at least one full purchase-consideration window, or you’re measuring the people who decide fast
  • Avoid running across a sale, a Black Friday, a school holiday or a bank holiday weekend, and if you can’t, say so at the top of the write-up
  • Anything longer than about six weeks is exposed to the site changing underneath it

When you can’t power it

Most UK ecommerce sites cannot power a classic A/B test on most of their pages, and pretending otherwise is the most common failure in the discipline. The ladder, in descending order of rigour:

  1. Test something bigger. Sample size scales with the square of the effect, so doubling the minimum detectable effect — the smallest change you’ve decided is worth finding — quarters the traffic needed. A 10% MDE at site level takes 1.1 months rather than 4.4. Be honest that you’re now testing a different question, and that “no detectable 10% effect” doesn’t rule out a real 4% one
  2. Move the metric closer to the change, as above
  3. Widen the exposure. Site-wide rather than one template, where the change makes sense in both
  4. Reduce the variance. Winsorisation and Capping on order value, and Variance Reduction where users have pre-period history
  5. Use a design built for continuous monitoring. Sequential Testing or Always-Valid Inference spend some power to buy the licence to stop early — mostly valuable for stopping on harm rather than on success
  6. Report Bayesian. Probability to Beat Control and a credible interval communicate an underpowered result more honestly than a p-value does, provided the prior is stated. It does not create information that isn’t there — Bayesian vs Frequentist
  7. Reframe as non-inferiority. When the real question is “can we ship this without damage”, Non-Inferiority Tests answer it with far less traffic than “is it better”
  8. Test demand before building. Painted Door Tests measure intent for a feature that doesn’t exist yet, at a fraction of the cost
  9. Ship it behind a flag with a holdout, and measure the accumulated effect of everything at once over a quarter — Holdout Groups · Feature Flags
  10. Decide on judgement, and say so. A documented judgement call is more honest than an underpowered test, because the test produces a number people will act on. Inconclusive Results

Test capacity is a property of your traffic, not your ambition. At 120,000 sessions a month with an average test needing ~1.5 months at full allocation, sequential running gives about eight tests a year. Getting to eighteen means running several at once on non-overlapping surfaces, which introduces Interaction Effects and needs the assignment to be independent — see Traffic Allocation. Plan the year against the arithmetic rather than against a target.

12. Build and QA

Before any traffic. Experiment QA owns the general procedure; the ecommerce-specific items:

  • Both variants render, on every device class, browser and viewport in your traffic, including the browsers you don’t personally use
  • The full purchase path completes in the variant, with a real transaction on staging and, if possible, one real transaction in production
  • Failure paths — declined card, out of stock mid-checkout, invalid address, expired session
  • Flicker. A client-side test that shows the original before swapping is measuring a worse experience than either variant, and it hurts the variant specifically — Flicker and Flash of Original Content · Client-Side vs Server-Side Testing
  • Speed impact of the test itself. A testing snippet in the critical path is a confound in every test you will ever run on that site
  • Assignment is recorded in analytics, so you can analyse outside the testing tool — Experiment Assignment Tracking
  • Assignment is stable across sessions, devices where identified, and page loads
  • Bots excluded from assignment as well as from analysis
  • Interaction with other running tests, personalisation rules, and any app or plugin that modifies the same surface
  • Consent interaction — does the test run for users who declined, are they assigned, and are they measured? An asymmetry here is a route to Sample Ratio Mismatch
  • Cached and CDN-served pages — that the variant survives the cache layer, and that the cache key includes the variant
  • A-A Tests if the machinery is new or hasn’t been checked in a year. It proves the pipeline before you trust it with a decision

13. While it runs

Don’t look at the result. A fixed-horizon test has a 5% false positive rate when checked once, at the end; every extra look is another chance to cross by luck. Peeking has the arithmetic, and the effect is large enough to change conclusions.

In plain terms: stopping the moment a test looks significant means you selected for the moment it looked good, and that moment arrives by chance in tests where nothing is happening.

Do check these daily, without looking at the primary metric:

  • Sample Ratio Mismatch — a 50/50 split arriving as 52/48 on large numbers invalidates the test whatever it says. First check, every day, no exceptions
  • Both variants still render. The most common failed experiment is one that quietly stopped showing the variant
  • Guardrails. Order errors, payment failures, page speed, JavaScript error rate, add-to-basket rate. Guardrail Metrics
  • Stock. A test whose winning variant promotes a product that sold out in week two is measuring stock, not design

Stop for harm, never for success. A pre-agreed harm threshold is the only legitimate early stop — Stopping Rules.

14. Read the result

Guide - Statistics for CRO has the ordered checks and the interval arithmetic. Run those first, then the four ecommerce-specific checks below, which are where the money is actually decided.

Check one: did contribution move

The trap, worked. A test lowers the free-delivery threshold from £60 to £40:

                              CONTROL      VARIANT

conversion rate                2.400%       2.640%    +10.0%  ← primary metric
goods AOV                      £62.00       £58.00     −6.5%
shipping revenue, blended       £1.10        £0.47   (12% of orders now pay)
                             ────────     ────────
total revenue per order        £63.10       £58.47

− cost of goods (45%)          £27.90       £26.10
− payment fees                 £ 1.40       £ 1.31
− pick and pack                £ 1.20       £ 1.20
− carriage                     £ 4.95       £ 4.95
                             ────────     ────────
contribution, kept             £27.65       £24.91
contribution, returned        −£11.05      −£10.96
blended at 12%                 £23.01       £20.60

REVENUE PER SESSION            £1.5144      £1.5436    +1.9%   ← also up
CONTRIBUTION PER SESSION       £0.5522      £0.5440    −1.5%   ← the truth
                                                    ≈ −£11,700 a year

Conversion rate up 10%, revenue up 1.9%, profit down. Every metric on a standard experiment dashboard reports a clear win. Only the contribution line catches it, and it catches it because the extra orders arrive smaller while carriage stays fixed at £4.95 whatever the basket is worth.

This is not an exotic case. It is the default outcome of any change that trades basket size for order count — thresholds, discount codes, entry-price product promotion, payment plans, and most bundle placements. See Shipping Thresholds and Discount Impact on Margin.

Check two: what happened to returns

Returns land 30–60 days after the order, which is usually after the test has been read and shipped. A change that removes friction from a decision people should have taken more slowly raises the return rate, and the £11.05 negative from the opening model does the rest:

same site, a variant that raises conversion 3% and returns 12% → 15%

contribution per order   0.85 × 27.65  +  0.15 × (−11.05)  =  £21.84   (was £23.01)
contribution per session          21.84 × 0.02472          =  £0.5400  (was £0.5522)

                                                              −2.2%
                                                              ≈ −£17,500 a year

So return rate is a guardrail that cannot be read inside the test window. Two workable responses: keep the assignment recorded and re-read the cohort at day 60, or refuse to ship anything that touches sizing, suitability, product information or returns policy without that lagged read. The first is better and requires only that assignment lives in the warehouse.

Check three: which customers, and what happens next

  • New versus returning. A win driven entirely by returning customers may be a change in timing rather than in demand
  • Category mix. Conversion up because cheaper stock got promoted is a merchandising change wearing a test’s clothes — Merchandising · Basket Composition
  • Cohort quality. Discount-acquired customers repeat at lower rates. A first-order win can be an LTV loss, and the test window can’t see it — Cohort Analysis · Customer Lifetime Value
  • Cannibalisation. Did the winning surface take orders from another surface rather than create them? Site-level conversion is the check

Check four: the standard statistical traps

Ranked by how often they actually bite in commercial testing:

Then state the result as a range in pounds, with a recommendation. Communicating Uncertainty and Reading a Test Result.

15. Ship, and check it landed

  • Ship behind a flag, so the winner goes live without a release and comes back off without one — Feature Flags · Progressive Delivery
  • Remove the test code. Accumulated dead variant code is a real and under-discussed cost on long-running programmes
  • Re-measure at 30 and 90 days. The shipped effect should be smaller than the test effect; if it’s absent, the test was a false positive or the implementation differs from the variant — Post-Test Validation
  • Re-read returns and cohort quality at 60 days, per §14
  • Archive the result, including losses and inconclusives, in a form someone can search in two years. Experiment Archive · Institutional Learning
  • Annotate the analytics, so next year’s trend investigation doesn’t rediscover this as an anomaly — Annotation and Change Logs
  • Update the funnel model. The baselines in §4 have moved, and every future prioritisation uses them

16. Make the programme compound

Individual tests barely matter. The system that produces them does.

The programme’s actual return, run for our example site at £795,000 annual contribution:

tests run in a year                                            18
win rate                                                      22%   → 4 winners
average validated lift of a winner                            +3%   → £23,900/yr each

gross value of winners                    4 × 23,900     =  £95,600
winner's curse and decay haircut          × 0.70         =  £66,900

losses prevented          14% of 18 = 2.5 tests at −2.4%
                          2.5 × 19,100                   =  £47,750
                          ← only counts if you would
                            genuinely have shipped them

programme cost            tooling, analysis time,
                          design and dev                 ≈ £55,000
                                                          ─────────
net, winners alone                                        +£11,900
net, including prevented losses                           +£59,650

The second line is the one that decides whether the programme is worth running, and it is the line nobody reports. Prevented losses are invisible by construction: nothing happened. Win Rate and Expected Value argues this at length, and it is the argument that keeps programmes funded through a bad quarter.

Verify the whole thing with a holdout. A 5% holdout excluded from every shipped change for a year is the only measurement of the programme’s real cumulative effect, and it routinely shows less than the sum of the individual test results. That gap is the honest number.

Then the things that raise the return more than any individual test:

  • Experimentation Velocity — tests per period is the dominant term in the arithmetic above, and it’s limited by traffic, dev capacity and decision latency rather than by ideas
  • Institutional Learning — turning results into beliefs about customers, and revising the beliefs when results contradict them. Without this you retest the same idea every eighteen months
  • Experimentation Maturity — identifying which stage is actually the bottleneck before investing in the others
  • Experiment Archive — searchable, including failures, with the hypothesis and the mechanism preserved rather than just the outcome

17. Conduct, law and the line

Not optional, and cheaper to consider now than after an enforcement letter.

[CHECK: current CMA position on drip pricing, urgency claims and review practices, and the operative consumer-protection regime, before relying on any specific statement of what is prohibited. This area has been actively legislated and enforced, and the detail matters more than any summary.]

18. The first ninety days

A defensible sequence when you take over an unfamiliar site.

Weeks 1–2 — trust the data. Access. Reconcile orders and revenue against finance. Purchase-event QA. Build the opening commercial model with finance’s real numbers. Output: a measurement findings list, and a model you can price things with.

Weeks 3–5 — see the shape. Funnel by segment. Headroom analysis. The rapid pass and the quantitative sweep. Speed audit. Three to five usability sessions on whatever the numbers located. Output: a prioritised backlog in pounds, and a written view of where the money is.

Weeks 6–8 — prove the machinery. An A/A test. Fix the outright defects found in weeks 3–5 without testing them. Ship the accessibility and legal items. Launch the first real test. Output: a working pipeline and something in flight.

Weeks 9–13 — establish the rhythm. Second and third tests. Set up the holdout. Write the archive template and the result write-up format. First quarterly report, including the losses prevented. Output: a programme rather than a project.

Ship the defects immediately and don’t wait for the testing programme. A quarter of measurement work with nothing shipped is how CRO engagements lose their sponsor in month two.

19. How programmes fail

Ordered by frequency, and each has a section above that prevents it.

  • Optimising conversion rate while contribution falls. §1 and §14
  • Testing on traffic that can’t detect anything, then acting on the numbers anyway. §11
  • Building the backlog from opinion rather than diagnosis. §6
  • Testing tiny changes because they’re easy to ship, which produces a year of underpowered nulls. §9
  • Never touching the commercial levers, because they belong to someone else. §8
  • Reading results early, and reading them charitably. §13
  • Not counting the losses prevented, so the programme looks like it isn’t working. §16
  • No archive, so the same test recurs every eighteen months. §15
  • Measurement never gated, so a quarter of results are artefacts. §3
  • Optimising the first order and ignoring the second. §1 and §14

The short version

  1. The objective is contribution per session. Conversion rate, revenue and AOV can all rise while it falls, and that is the normal outcome of threshold, discount and bundling changes
  2. Prove the numbers before analysing them, and expect a measurable share of “conversion problems” to be tracking problems
  3. Drop-off size doesn’t indicate value. In a multiplicative funnel a 10% improvement is worth the same at every step; what differs is headroom, achievability and detectability
  4. Calculate the sample size before designing anything. Most tests are unaffordable and it takes ten minutes to find out
  5. Measure as close to the change as the question allows. The same effect can be 1.7× faster to detect on a deeper metric
  6. A returned order costs money, and the return arrives after the test is read
  7. Don’t test what you’d ship anyway, and don’t test outright defects. Fix them
  8. Count the losses you prevented. It’s half the programme’s value and none of its reporting

Related, in the order you’d reach for them: Guide - Statistics for CRO for the arithmetic of a single test · Guide - Running an Experiment for the procedure around it · Guide - Running a Conversion Audit for the diagnostic sweep in §6 at full depth · Guide - Unit Economics and Commercial Decisions for the model at the top · Guide - Diagnosing and Fixing Performance for §8’s speed lever. Full cross-domain view: CRO.