Tags: experimentation ux guide

Guide - Running a Conversion Audit

Date: 2026-09-28


An audit turns a site you don’t yet understand into a ranked list of findings someone can act on. A finding needs three things: where the loss is, why it happens, and how sure you are. Most audits deliver the first and dress the other two up with screenshots. The whole procedure is getting every row to that standard, then sorting the rows into fix, test, research or leave.


A conversion audit is a time-boxed diagnostic of where and why a site loses customers it could have kept, ending in an evidenced, prioritised backlog.

It is the diagnostic half of Guide - Ecommerce CRO (§6 there) done at full depth, as a bounded piece of work. That guide owns the commercial model, access, the measurement gate, the surface checklist and the expected-value arithmetic; this one links to them by section rather than repeating them.

Same invented retailer throughout, so every figure checks against the other guide: UK direct-to-consumer, 120,000 sessions a month, 2.40% conversion, £23.01 contribution per order after returns, funnel as in Guide - Ecommerce CRO §4.

Written for an in-house audit. §10 covers what changes on a client’s site, and §11 what changes when the site sells leads or subscriptions rather than products.

The shape of it

STAGE              QUESTION                          PRODUCES                     TIME*
1 Frame            what decision does this serve?    brief, access, time box      ½ day
2 Gate             are the numbers real?             known limits                 1–2 days
3 Walk it cold     what does a stranger hit?         hypotheses  (H-rows)         1 day
4 Locate           where is the loss, in £?          leak table  (L-rows)         2–3 days
5 Explain          why, at the located steps only?   evidence against rows        4–6 days
6 Triangulate      what do we believe, how firmly?   graded findings              1 day
7 Sort             fix, test, research, or leave?    backlog                      1 day
8 Write up         what does the reader do Monday?   report                       1–2 days

                                                     * a three-week audit, one person

Everything writes to one table — the finding register, below. Stages 3 and 4 add rows; stages 5 and 6 fill in columns; stage 7 reads it. The report at the end is a view of the register, not a separate piece of writing.

The finding register

The artefact the whole audit exists to produce. Copy it into a sheet before starting.

ID   SOURCE  WHERE                  WHAT HAPPENS                    SIZE (£/yr,    EVIDENCE                     MECHANISM                   GRADE  ACTION
                                                                    upper bound)
F1   L+H     checkout › delivery    postcode rejected, re-entered   94,000 ┐       form analytics: 2.3 re-      valid formats rejected;     A      FIX validation
             step, mobile-new       repeatedly, then exit                  │       entries per starter;         error doesn't say what             TEST lookup
                                                                           │       replay 14 of 30; H: error    format it wants
                                                                           │       prevention
F2   L       search, all devices    9% of searches return nothing;  17,700 │       search logs; path            vocabulary mismatch —       B      FIX synonyms
                                    those sessions convert 1.2%            │       analysis: exit or loop       customers' words not the
                                    vs 4.5%                                │       to home                      catalogue's
F3   L+V     checkout › delivery    delivery cost first seen at     76,000 ┘       funnel: 9% vs 4% exit;       cost surprise at the        B      TEST delivery cost
             step, all              checkout                                shared VoC: "how much is       point of commitment                on product page
                                                                        step     delivery" top-3 ticket
F4   L       checkout › payment,    wallet button fails inside      12,700        error logs: 1.1% of          defect                      A      FIX
             in-app browsers        social in-app browsers                         starts; 60% of those exit
F5   L       product page, mobile   slow at p75; slowest quartile   unsized       field data; correlation      unknown — speed or the      C      RESEARCH speed
                                    converts worst                                 only                         traffic that's slow?               holdout
F6   H       product page, mobile   long titles push add-to-basket  unsized       heuristic only; no replay    —                           D      LEAVE, logged
                                    below the fold                                 or funnel signal

SOURCE   H heuristic walk · L located in data · V voice of customer

Three things to notice before going further:

  • F1 and F3 share a step, so their sizes don’t add. Both draw on the same mobile checkout loss. Never sum the size column — it double-counts every step with more than one finding, and the total is the most-quoted and least-true number in audit reports
  • Sizes are upper bounds. Each is the gap to a better segment, and part of every such gap is the population rather than the page — Guide - Ecommerce CRO §5 and Confounding Variables
  • Grade and action are separate columns. F1 and F4 are both grade A and neither gets a test, because both are defects. Confidence decides whether you know enough; the nature of the mechanism decides what to do about it

1. Frame it

  • The decision the audit serves. “Where should the next two quarters of CRO effort go?” or “why did mobile conversion fall after the replatform?” produce different audits. “Look at the site” produces a list
  • The time box. Three weeks is enough for a mid-size ecommerce site with one auditor. Set it before starting, or stage 5 expands until the deadline
  • Scope. Templates, markets, devices. Name what’s out of scope so it’s a decision, not an omission
  • Access — Guide - Ecommerce CRO §2 has the list. Chase it on day one; missing replay or search logs changes which methods are available in stage 5
  • The one person who can say yes to a price, delivery or returns change. Many of the highest-value findings are commercial rather than interface, and without this person they go straight to “leave”

2. Gate: are the numbers real

Run Guide - Auditing a Tracking Plan if the site is new to you, or at least the ecommerce minimum in Guide - Ecommerce CRO §3: purchase reconciliation, revenue units, duplicates, purchase-on-failure, bots, consent coverage, checkout observability.

The output is a list of known limits, not a clean bill of health. “Hosted checkout steps 2–3 unobservable”, “consent coverage 71% on mobile, 84% on desktop” — each goes at the top of the register as a caveat on every row it touches. A device gap in consent coverage contaminates every device comparison in stage 4, so find it now.

3. Walk it cold — before the data

Do this before opening analytics. Once you’ve seen that mobile checkout is the leak, you’ll find mobile checkout problems whether they’re there or not. The walk-through is the only stage that can be done without that bias, and only once.

The rapid pass — Guide - Ecommerce CRO §6 has the list. Buy something on a phone on mobile data as a new customer; search using a customer’s word; try to find the returns policy; tab through checkout.

Then a structured review against two sets at once — see Heuristic Evaluation:

  • a usability set (Nielsen’s ten) — where people can’t proceed
  • a persuasion set — value proposition, relevance, clarity, anxiety, distraction — where they don’t want to. See Value Propositions, Trust Signals and Risk Reversal

Walk the templates in the order traffic hits them, one pass for flow and one for detail.

Log everything as H-rows: hypotheses, not findings. One evaluator finds about a third of the problems a panel would, and flags some that no user ever trips over. The rows earn their place only if stage 4 or 5 corroborates them. For the critical path, borrow a second evaluator for an hour and don’t compare lists until both are done.

Also note on the walk: accessibility failures. They’re fixed rather than tested and carry legal weight — Accessibility Law in the UK and WCAG.

4. Locate the loss, in pounds

Build the funnel by segment — the method is Guide - Ecommerce CRO §4, and the headroom logic is §5. What the audit adds is discipline about what counts as located:

A location is a step × a segment × a size. “Checkout is bad” isn’t located. “Delivery step, mobile, new customers, 31% completion against 48% for desktop-new, upper bound £94,000 a year” is.

The sweep, cheapest first, each adding L-rows or corroborating H-rows:

Sizing each row, worked for F2:

sessions using search           120,000 × 18%            =  21,600 / month
searches returning nothing      21,600 × 9%              =   1,944 sessions
conversion, results found       4.5%
conversion, zero results        1.2%
upper bound: all zero-result sessions convert like the rest
                                1,944 × (0.045 − 0.012)  =      64 orders / month
                                64 × £23.01 × 12         ≈  £17,700 / year

Upper bound twice over: some zero-result searches are for things the shop doesn’t sell, and no synonym list converts those. The true recoverable share is unknown until the fix ships.

Rows with no size are allowed — F5 and F6 above — but they rank below every sized row until something sizes them.

5. Explain — at the located steps only

Qualitative effort goes where stage 4 pointed, and nowhere else. Watching replays of the whole site is entertainment; watching thirty of the located step, filtered to the located segment, is evidence. See Qualitative vs Quantitative Research.

Each method, with how much is enough and what it can’t tell you:

MethodHow muchWhat it producesWhat it can’t tell you
Session Replay20–40 sessions of the step × segment; stop when five in a row add no new codethe observed behaviour, coded and countedwhy; and it’s a biased sample — replay tools sample, and consent and blockers remove visitors
Form Analyticsthe whole population, if instrumentedfield-level abandonment, re-entry, time, error ratewhether the field is the cause or the moment people gave up
Heatmapsone per located template × deviceaggregate attention and interactionanything on dynamic or personalised pages; causes
Voice of Customer Data3–6 months of support tickets, chat, reviews, returns reasonsthe customer’s own vocabulary for the problem, and its frequencythe silent majority who never write in
Surveyson-step exit survey, one open question; post-purchase “what nearly stopped you?”; 100–200 responses to codestated reasons, at volumebehaviour — what people say and do diverge
Usability Testingfive participants per distinct user group, on the located taskswhy, fastest of any methodfrequency — five sessions show a problem exists, not how common it is. Sample Size in Qualitative Research

The replay tagging sheet — the difference between “I watched some replays” and evidence:

REPLAY LOG   step: checkout › delivery   segment: mobile, new   filter: reached step, did not complete

#    session   reached      codes                        note
1    a81f…     delivery     POSTCODE-REJECT, RETRY×3     UK format with no space
2    c02e…     delivery     COST-SEEN, EXIT              exits 4s after delivery price renders
3    9b7d…     delivery     POSTCODE-REJECT, EXIT
4    e11a…     delivery     FIELD-FOCUS-LOOP             keyboard covers the error message
…
30   77c3…     delivery     COST-SEEN, BACK, EXIT

TALLY    POSTCODE-REJECT 14/30 · COST-SEEN→EXIT 9/30 · KEYBOARD-OCCLUDES 5/30 · unexplained 6/30

Codes emerge from the first ten sessions and then stay fixed, so the tally counts something consistent — Qualitative Coding. Keep the unexplained count. Six of thirty unexplained is a finding too: something else is going on, and the report shouldn’t pretend otherwise.

For choosing between methods in more depth: Guide - User Research Methods.

6. Triangulate and grade

Two sources of different kinds agreeing is worth more than any amount of one source. Replay and form analytics both saying “postcode” is corroboration; ten more replays saying it is not. Triangulation owns the discipline; the audit applies it as a grade per row:

GRADE  STANDARD                                                     USUAL ACTION
A      located in data, AND the mechanism shown by an independent   fix if a defect, otherwise test
       second source of a different kind
B      located, mechanism supported by one qualitative source       test
C      located, mechanism unknown                                   research first
D      not located — heuristic, opinion or anecdote only             leave, logged

When sources disagree, they’re usually both right about different people. Replay shows mobile users stalling at postcode; the survey says delivery cost. Check whether they’re the same segment before choosing — here the survey fires on all devices and replay was filtered to mobile, so F1 and F3 are two findings, not one contested one.

Grade honestly downwards. A row resting on your own heuristic review plus one replay is a B at best, however obvious it feels. The grade is what makes the report trustworthy, and one inflated A that fails its test costs the credibility of every other row.

In plain terms: the grade says how much of what you believe about a finding came from evidence rather than from you.

7. Sort into fix, test, research, leave

The decision most audits get wrong in both directions — testing defects, and shipping opinions untested.

Is the mechanism a defect — broken, erroring, rejecting valid input, inaccessible?
  └─ yes → FIX. No test. Verify after release with the metric that located it

Is the change cheap, obviously beneficial and something you'd ship anyway?
  └─ yes → FIX, or a non-inferiority test if there's any risk it hurts

Grade A or B, and can the step detect a plausible effect in acceptable time?
  └─ yes → TEST — hypothesis written from the row
  └─ no  → ship with a before/after read and say plainly it's unproven,
           or move the test to a higher-traffic proxy metric

Grade C?
  └─ RESEARCH — name the method that would raise it to B, and its cost

Grade D?
  └─ LEAVE — but keep it in the register. Next audit's replay may locate it

A TEST row becomes a hypothesis with the row’s evidence and mechanism in it — Hypothesis Design. Detectability is Minimum Detectable Effect against the step’s traffic — Guide - Ecommerce CRO §11 has the arithmetic. It’s why deep-funnel findings often go to “ship and read” on smaller sites.

Then rank within each bucket:

  • Fixes by size ÷ effort. They don’t need a win rate — defects are fixed with certainty
  • Tests by expected value, Guide - Ecommerce CRO §9. The grade feeds P(it works): an A-row starts near the top of your programme’s win rate, a B-row nearer the bottom — Win Rate and Expected Value
                     upper bound   P(works)   decay    expected value
F1 address lookup    £94,000       0.35       × 0.7    £23,000
F3 delivery cost     £76,000       0.25       × 0.7    £13,300
                     ↑ same step — run sequentially, not in parallel, and re-size F3
                       after F1 ships, because F1 changes the population F3 sees

Look at the list for what’s missing before handing it over: speed, delivery proposition, payment methods, accessibility, and the mobile version of each. A backlog built from diagnosis skews toward what’s visible in replay.

8. Write it up

The reader wants to know what to do on Monday, not what you did for three weeks. Method goes at the back.

REPORT SKELETON

1  The answer                  one paragraph: where the money is, the three things
                               to do first, the one thing to stop doing
2  What we can't see           known limits from stage 2, stated before any number —
                               "checkout steps 2–3 unobservable"
3  Fixes                       F-rows marked FIX, ranked. Each: what, where, evidence,
                               size, how you'll know it worked
4  Tests                       ranked by expected value. Each: the hypothesis,
                               the metric, the traffic it needs, how long
5  Research                    C-rows, the method that would resolve each, its cost
6  Parked                      D-rows, one line each — so nobody re-finds them
7  Method and sources          what was run, how much, dates, the replay log,
                               the full register as an appendix

Rules for the numbers in it:

  • Every size is labelled as an upper bound, and none are summed — Communicating Uncertainty
  • Give a range where you have one. “£40,000–£94,000” with the reason for the lower end beats “£94,000” every time it’s checked later
  • One screenshot per finding, annotated with the thing to notice. Twenty screenshots per finding is a portfolio, not a report
  • Put the grade next to the claim, in the text, not in an appendix

Store the register where the next audit will find it — Research Repositories. The D-rows from this audit are the cheapest hypotheses the next one will have.

9. How audits fail

  • Data first, walk second. The walk-through then confirms the data instead of generating independent hypotheses
  • Summing the size column. Findings on the same step aren’t additive, and the upper bounds were never achievable together
  • Findings without locations. A heuristic list presented as the audit — every row a D, formatted like an A
  • Replay as browsing. Unfiltered watching produces vivid anecdotes, and the vivid one leads the report
  • Testing defects. Two months of traffic spent proving that a broken postcode field is worse than a working one
  • A “best practice” section. Advice that isn’t located on this site is a D-row by definition
  • Skipping the gate. A quarter of “findings” turning out to be tracking — Guide - Auditing a Tracking Plan
  • No owner for commercial findings. Delivery, pricing and returns rows go nowhere without the person from §1
  • Stopping at the report. The audit’s value is realised in the first two fixes shipping. Book the follow-up read before handing over

10. When it’s a client’s site

The procedure doesn’t change. What changes:

  • Access takes longer and arrives partially. Scope the methods to the access that’s confirmed, not promised, and say in the report which ones were unavailable
  • The report is the product. Stage 8 gets twice the time, and §1 of the report is written for someone who won’t read §3
  • Their data, their history. Ask what’s been tested before. Re-proposing a test they ran last year costs more credibility than any finding earns
  • Known limits are awkward to deliver. “Your tracking double-counts purchases on refresh” is unwelcome — which is why it goes in section 2 of the report, stated neutrally, before anyone is invested in the numbers that depend on it
  • Anonymise before anything leaves the engagement — screenshots included

11. When it’s not ecommerce

The procedure is the same; the funnel ends somewhere else.

  • Lead generation. The form is the checkout, and the conversion that matters is a qualified lead, which happens days later in a CRM. Locate on qualified outcomes where you can join them — Lead Funnel Stages, Offline Conversion Imports. An audit optimising form completions alone will happily recommend removing the fields that qualified people — Lead Quality vs Volume. Response time belongs in scope too — Speed to Lead
  • Subscription and SaaS. Sign-up is the middle of the funnel, not the end. Locate on activation and first-month retention, not on sign-ups — Activation and Time to Value. The walk-through covers onboarding, and the voice-of-customer sweep includes cancellation reasons — Cancellation Flows
  • Size per qualified outcome, not per form fill. Contribution per order becomes value per qualified lead or per retained subscriber, and every size in the register uses it

The short version

  1. Frame the decision and the time box; get access and the commercial owner on day one
  2. Gate the data; write known limits at the top of the register
  3. Walk it cold before opening analytics; log hypotheses, not findings
  4. Locate: step × segment × £ upper bound. Never sum the column
  5. Explain only the located steps; code the replays and keep the unexplained count
  6. Grade every row by evidence, honestly downwards
  7. Defects are fixed, not tested; A/B-grade rows are tested by expected value; C-rows get research; D-rows are parked, not deleted
  8. Report what to do Monday, then what you can’t see, then the evidence

Related: Guide - Ecommerce CRO for the commercial model and arithmetic this sits inside · Guide - Running an Experiment for what happens to the TEST rows · Guide - Statistics for CRO for reading them · Guide - Diagnosing and Fixing Performance for rows like F5. Full cross-domain view: CRO.