Tags: experimentation ux guide
Guide - Running a Conversion Audit
Date: 2026-09-28
An audit turns a site you don’t yet understand into a ranked list of findings someone can act on. A finding needs three things: where the loss is, why it happens, and how sure you are. Most audits deliver the first and dress the other two up with screenshots. The whole procedure is getting every row to that standard, then sorting the rows into fix, test, research or leave.
A conversion audit is a time-boxed diagnostic of where and why a site loses customers it could have kept, ending in an evidenced, prioritised backlog.
It is the diagnostic half of Guide - Ecommerce CRO (§6 there) done at full depth, as a bounded piece of work. That guide owns the commercial model, access, the measurement gate, the surface checklist and the expected-value arithmetic; this one links to them by section rather than repeating them.
Same invented retailer throughout, so every figure checks against the other guide: UK direct-to-consumer, 120,000 sessions a month, 2.40% conversion, £23.01 contribution per order after returns, funnel as in Guide - Ecommerce CRO §4.
Written for an in-house audit. §10 covers what changes on a client’s site, and §11 what changes when the site sells leads or subscriptions rather than products.
The shape of it
STAGE QUESTION PRODUCES TIME*
1 Frame what decision does this serve? brief, access, time box ½ day
2 Gate are the numbers real? known limits 1–2 days
3 Walk it cold what does a stranger hit? hypotheses (H-rows) 1 day
4 Locate where is the loss, in £? leak table (L-rows) 2–3 days
5 Explain why, at the located steps only? evidence against rows 4–6 days
6 Triangulate what do we believe, how firmly? graded findings 1 day
7 Sort fix, test, research, or leave? backlog 1 day
8 Write up what does the reader do Monday? report 1–2 days
* a three-week audit, one person
Everything writes to one table — the finding register, below. Stages 3 and 4 add rows; stages 5 and 6 fill in columns; stage 7 reads it. The report at the end is a view of the register, not a separate piece of writing.
The finding register
The artefact the whole audit exists to produce. Copy it into a sheet before starting.
ID SOURCE WHERE WHAT HAPPENS SIZE (£/yr, EVIDENCE MECHANISM GRADE ACTION
upper bound)
F1 L+H checkout › delivery postcode rejected, re-entered 94,000 ┐ form analytics: 2.3 re- valid formats rejected; A FIX validation
step, mobile-new repeatedly, then exit │ entries per starter; error doesn't say what TEST lookup
│ replay 14 of 30; H: error format it wants
│ prevention
F2 L search, all devices 9% of searches return nothing; 17,700 │ search logs; path vocabulary mismatch — B FIX synonyms
those sessions convert 1.2% │ analysis: exit or loop customers' words not the
vs 4.5% │ to home catalogue's
F3 L+V checkout › delivery delivery cost first seen at 76,000 ┘ funnel: 9% vs 4% exit; cost surprise at the B TEST delivery cost
step, all checkout shared VoC: "how much is point of commitment on product page
step delivery" top-3 ticket
F4 L checkout › payment, wallet button fails inside 12,700 error logs: 1.1% of defect A FIX
in-app browsers social in-app browsers starts; 60% of those exit
F5 L product page, mobile slow at p75; slowest quartile unsized field data; correlation unknown — speed or the C RESEARCH speed
converts worst only traffic that's slow? holdout
F6 H product page, mobile long titles push add-to-basket unsized heuristic only; no replay — D LEAVE, logged
below the fold or funnel signal
SOURCE H heuristic walk · L located in data · V voice of customer
Three things to notice before going further:
- F1 and F3 share a step, so their sizes don’t add. Both draw on the same mobile checkout loss. Never sum the size column — it double-counts every step with more than one finding, and the total is the most-quoted and least-true number in audit reports
- Sizes are upper bounds. Each is the gap to a better segment, and part of every such gap is the population rather than the page — Guide - Ecommerce CRO §5 and Confounding Variables
- Grade and action are separate columns. F1 and F4 are both grade A and neither gets a test, because both are defects. Confidence decides whether you know enough; the nature of the mechanism decides what to do about it
1. Frame it
- The decision the audit serves. “Where should the next two quarters of CRO effort go?” or “why did mobile conversion fall after the replatform?” produce different audits. “Look at the site” produces a list
- The time box. Three weeks is enough for a mid-size ecommerce site with one auditor. Set it before starting, or stage 5 expands until the deadline
- Scope. Templates, markets, devices. Name what’s out of scope so it’s a decision, not an omission
- Access — Guide - Ecommerce CRO §2 has the list. Chase it on day one; missing replay or search logs changes which methods are available in stage 5
- The one person who can say yes to a price, delivery or returns change. Many of the highest-value findings are commercial rather than interface, and without this person they go straight to “leave”
2. Gate: are the numbers real
Run Guide - Auditing a Tracking Plan if the site is new to you, or at least the ecommerce minimum in Guide - Ecommerce CRO §3: purchase reconciliation, revenue units, duplicates, purchase-on-failure, bots, consent coverage, checkout observability.
The output is a list of known limits, not a clean bill of health. “Hosted checkout steps 2–3 unobservable”, “consent coverage 71% on mobile, 84% on desktop” — each goes at the top of the register as a caveat on every row it touches. A device gap in consent coverage contaminates every device comparison in stage 4, so find it now.
3. Walk it cold — before the data
Do this before opening analytics. Once you’ve seen that mobile checkout is the leak, you’ll find mobile checkout problems whether they’re there or not. The walk-through is the only stage that can be done without that bias, and only once.
The rapid pass — Guide - Ecommerce CRO §6 has the list. Buy something on a phone on mobile data as a new customer; search using a customer’s word; try to find the returns policy; tab through checkout.
Then a structured review against two sets at once — see Heuristic Evaluation:
- a usability set (Nielsen’s ten) — where people can’t proceed
- a persuasion set — value proposition, relevance, clarity, anxiety, distraction — where they don’t want to. See Value Propositions, Trust Signals and Risk Reversal
Walk the templates in the order traffic hits them, one pass for flow and one for detail.
Log everything as H-rows: hypotheses, not findings. One evaluator finds about a third of the problems a panel would, and flags some that no user ever trips over. The rows earn their place only if stage 4 or 5 corroborates them. For the critical path, borrow a second evaluator for an hour and don’t compare lists until both are done.
Also note on the walk: accessibility failures. They’re fixed rather than tested and carry legal weight — Accessibility Law in the UK and WCAG.
4. Locate the loss, in pounds
Build the funnel by segment — the method is Guide - Ecommerce CRO §4, and the headroom logic is §5. What the audit adds is discipline about what counts as located:
A location is a step × a segment × a size. “Checkout is bad” isn’t located. “Delivery step, mobile, new customers, 31% completion against 48% for desktop-new, upper bound £94,000 a year” is.
The sweep, cheapest first, each adding L-rows or corroborating H-rows:
- Funnel Analysis by device, new/returning, channel, landing template — Segmentation (analysis)
- Form Analytics — names the field, not just the step. On checkout, the highest-yield single source
- Internal search logs — zero-result rate, exits after search, the words used — Search and Findability
- Path Analysis — loops back to search or home are navigation failures
- Error logs — validation failures by field, payment declines by method, JavaScript errors by browser — Error Tracking
- Speed by template and device at p75 from field data — Core Web Vitals
- Stock — a conversion “problem” that is actually Stockouts and Availability
Sizing each row, worked for F2:
sessions using search 120,000 × 18% = 21,600 / month
searches returning nothing 21,600 × 9% = 1,944 sessions
conversion, results found 4.5%
conversion, zero results 1.2%
upper bound: all zero-result sessions convert like the rest
1,944 × (0.045 − 0.012) = 64 orders / month
64 × £23.01 × 12 ≈ £17,700 / year
Upper bound twice over: some zero-result searches are for things the shop doesn’t sell, and no synonym list converts those. The true recoverable share is unknown until the fix ships.
Rows with no size are allowed — F5 and F6 above — but they rank below every sized row until something sizes them.
5. Explain — at the located steps only
Qualitative effort goes where stage 4 pointed, and nowhere else. Watching replays of the whole site is entertainment; watching thirty of the located step, filtered to the located segment, is evidence. See Qualitative vs Quantitative Research.
Each method, with how much is enough and what it can’t tell you:
| Method | How much | What it produces | What it can’t tell you |
|---|---|---|---|
| Session Replay | 20–40 sessions of the step × segment; stop when five in a row add no new code | the observed behaviour, coded and counted | why; and it’s a biased sample — replay tools sample, and consent and blockers remove visitors |
| Form Analytics | the whole population, if instrumented | field-level abandonment, re-entry, time, error rate | whether the field is the cause or the moment people gave up |
| Heatmaps | one per located template × device | aggregate attention and interaction | anything on dynamic or personalised pages; causes |
| Voice of Customer Data | 3–6 months of support tickets, chat, reviews, returns reasons | the customer’s own vocabulary for the problem, and its frequency | the silent majority who never write in |
| Surveys | on-step exit survey, one open question; post-purchase “what nearly stopped you?”; 100–200 responses to code | stated reasons, at volume | behaviour — what people say and do diverge |
| Usability Testing | five participants per distinct user group, on the located tasks | why, fastest of any method | frequency — five sessions show a problem exists, not how common it is. Sample Size in Qualitative Research |
The replay tagging sheet — the difference between “I watched some replays” and evidence:
REPLAY LOG step: checkout › delivery segment: mobile, new filter: reached step, did not complete
# session reached codes note
1 a81f… delivery POSTCODE-REJECT, RETRY×3 UK format with no space
2 c02e… delivery COST-SEEN, EXIT exits 4s after delivery price renders
3 9b7d… delivery POSTCODE-REJECT, EXIT
4 e11a… delivery FIELD-FOCUS-LOOP keyboard covers the error message
…
30 77c3… delivery COST-SEEN, BACK, EXIT
TALLY POSTCODE-REJECT 14/30 · COST-SEEN→EXIT 9/30 · KEYBOARD-OCCLUDES 5/30 · unexplained 6/30
Codes emerge from the first ten sessions and then stay fixed, so the tally counts something consistent — Qualitative Coding. Keep the unexplained count. Six of thirty unexplained is a finding too: something else is going on, and the report shouldn’t pretend otherwise.
For choosing between methods in more depth: Guide - User Research Methods.
6. Triangulate and grade
Two sources of different kinds agreeing is worth more than any amount of one source. Replay and form analytics both saying “postcode” is corroboration; ten more replays saying it is not. Triangulation owns the discipline; the audit applies it as a grade per row:
GRADE STANDARD USUAL ACTION
A located in data, AND the mechanism shown by an independent fix if a defect, otherwise test
second source of a different kind
B located, mechanism supported by one qualitative source test
C located, mechanism unknown research first
D not located — heuristic, opinion or anecdote only leave, logged
When sources disagree, they’re usually both right about different people. Replay shows mobile users stalling at postcode; the survey says delivery cost. Check whether they’re the same segment before choosing — here the survey fires on all devices and replay was filtered to mobile, so F1 and F3 are two findings, not one contested one.
Grade honestly downwards. A row resting on your own heuristic review plus one replay is a B at best, however obvious it feels. The grade is what makes the report trustworthy, and one inflated A that fails its test costs the credibility of every other row.
In plain terms: the grade says how much of what you believe about a finding came from evidence rather than from you.
7. Sort into fix, test, research, leave
The decision most audits get wrong in both directions — testing defects, and shipping opinions untested.
Is the mechanism a defect — broken, erroring, rejecting valid input, inaccessible?
└─ yes → FIX. No test. Verify after release with the metric that located it
Is the change cheap, obviously beneficial and something you'd ship anyway?
└─ yes → FIX, or a non-inferiority test if there's any risk it hurts
Grade A or B, and can the step detect a plausible effect in acceptable time?
└─ yes → TEST — hypothesis written from the row
└─ no → ship with a before/after read and say plainly it's unproven,
or move the test to a higher-traffic proxy metric
Grade C?
└─ RESEARCH — name the method that would raise it to B, and its cost
Grade D?
└─ LEAVE — but keep it in the register. Next audit's replay may locate it
A TEST row becomes a hypothesis with the row’s evidence and mechanism in it — Hypothesis Design. Detectability is Minimum Detectable Effect against the step’s traffic — Guide - Ecommerce CRO §11 has the arithmetic. It’s why deep-funnel findings often go to “ship and read” on smaller sites.
Then rank within each bucket:
- Fixes by size ÷ effort. They don’t need a win rate — defects are fixed with certainty
- Tests by expected value, Guide - Ecommerce CRO §9. The grade feeds P(it works): an A-row starts near the top of your programme’s win rate, a B-row nearer the bottom — Win Rate and Expected Value
upper bound P(works) decay expected value
F1 address lookup £94,000 0.35 × 0.7 £23,000
F3 delivery cost £76,000 0.25 × 0.7 £13,300
↑ same step — run sequentially, not in parallel, and re-size F3
after F1 ships, because F1 changes the population F3 sees
Look at the list for what’s missing before handing it over: speed, delivery proposition, payment methods, accessibility, and the mobile version of each. A backlog built from diagnosis skews toward what’s visible in replay.
8. Write it up
The reader wants to know what to do on Monday, not what you did for three weeks. Method goes at the back.
REPORT SKELETON
1 The answer one paragraph: where the money is, the three things
to do first, the one thing to stop doing
2 What we can't see known limits from stage 2, stated before any number —
"checkout steps 2–3 unobservable"
3 Fixes F-rows marked FIX, ranked. Each: what, where, evidence,
size, how you'll know it worked
4 Tests ranked by expected value. Each: the hypothesis,
the metric, the traffic it needs, how long
5 Research C-rows, the method that would resolve each, its cost
6 Parked D-rows, one line each — so nobody re-finds them
7 Method and sources what was run, how much, dates, the replay log,
the full register as an appendix
Rules for the numbers in it:
- Every size is labelled as an upper bound, and none are summed — Communicating Uncertainty
- Give a range where you have one. “£40,000–£94,000” with the reason for the lower end beats “£94,000” every time it’s checked later
- One screenshot per finding, annotated with the thing to notice. Twenty screenshots per finding is a portfolio, not a report
- Put the grade next to the claim, in the text, not in an appendix
Store the register where the next audit will find it — Research Repositories. The D-rows from this audit are the cheapest hypotheses the next one will have.
9. How audits fail
- Data first, walk second. The walk-through then confirms the data instead of generating independent hypotheses
- Summing the size column. Findings on the same step aren’t additive, and the upper bounds were never achievable together
- Findings without locations. A heuristic list presented as the audit — every row a D, formatted like an A
- Replay as browsing. Unfiltered watching produces vivid anecdotes, and the vivid one leads the report
- Testing defects. Two months of traffic spent proving that a broken postcode field is worse than a working one
- A “best practice” section. Advice that isn’t located on this site is a D-row by definition
- Skipping the gate. A quarter of “findings” turning out to be tracking — Guide - Auditing a Tracking Plan
- No owner for commercial findings. Delivery, pricing and returns rows go nowhere without the person from §1
- Stopping at the report. The audit’s value is realised in the first two fixes shipping. Book the follow-up read before handing over
10. When it’s a client’s site
The procedure doesn’t change. What changes:
- Access takes longer and arrives partially. Scope the methods to the access that’s confirmed, not promised, and say in the report which ones were unavailable
- The report is the product. Stage 8 gets twice the time, and §1 of the report is written for someone who won’t read §3
- Their data, their history. Ask what’s been tested before. Re-proposing a test they ran last year costs more credibility than any finding earns
- Known limits are awkward to deliver. “Your tracking double-counts purchases on refresh” is unwelcome — which is why it goes in section 2 of the report, stated neutrally, before anyone is invested in the numbers that depend on it
- Anonymise before anything leaves the engagement — screenshots included
11. When it’s not ecommerce
The procedure is the same; the funnel ends somewhere else.
- Lead generation. The form is the checkout, and the conversion that matters is a qualified lead, which happens days later in a CRM. Locate on qualified outcomes where you can join them — Lead Funnel Stages, Offline Conversion Imports. An audit optimising form completions alone will happily recommend removing the fields that qualified people — Lead Quality vs Volume. Response time belongs in scope too — Speed to Lead
- Subscription and SaaS. Sign-up is the middle of the funnel, not the end. Locate on activation and first-month retention, not on sign-ups — Activation and Time to Value. The walk-through covers onboarding, and the voice-of-customer sweep includes cancellation reasons — Cancellation Flows
- Size per qualified outcome, not per form fill. Contribution per order becomes value per qualified lead or per retained subscriber, and every size in the register uses it
The short version
- Frame the decision and the time box; get access and the commercial owner on day one
- Gate the data; write known limits at the top of the register
- Walk it cold before opening analytics; log hypotheses, not findings
- Locate: step × segment × £ upper bound. Never sum the column
- Explain only the located steps; code the replays and keep the unexplained count
- Grade every row by evidence, honestly downwards
- Defects are fixed, not tested; A/B-grade rows are tested by expected value; C-rows get research; D-rows are parked, not deleted
- Report what to do Monday, then what you can’t see, then the evidence
Related: Guide - Ecommerce CRO for the commercial model and arithmetic this sits inside · Guide - Running an Experiment for what happens to the TEST rows · Guide - Statistics for CRO for reading them · Guide - Diagnosing and Fixing Performance for rows like F5. Full cross-domain view: CRO.