Tags: analytics concept

Benchmarking

Date: 2026-08-17


Comparing your numbers against industry figures. Nearly always invalid, because the published figure is a differently-defined metric measured on a differently-composed population by a party with an interest in the answer — and the comparison gets made anyway, usually in a board meeting.


Benchmarking is judging a metric against an external reference figure — an industry average, a competitor, a vendor report — rather than against your own history.

The four reasons it doesn’t hold

1. Definitions differ, silently.

"average ecommerce conversion rate: 2.6%"

conversion of WHAT?
  sessions?  users?  new users only?
  including bots?  including internal traffic?
  is a subscription renewal a conversion?
  is a click to a marketplace listing?

your 1.9% and their 2.6% may be the same performance
measured two ways — or a real 40% gap. nothing in the
published figure tells you which.

Different denominators alone can account for the entire apparent gap. Sessions-based and user-based conversion rates differ by roughly the sessions-per-user ratio, which is typically 1.3–2.0 — enough to move 1.9% to 3.0% with no change in behaviour — Metric Design, User Counting.

2. The population differs. “Ecommerce” spans £8 impulse purchases and £900 considered ones. Traffic mix matters more than category: a site that’s 70% branded search will out-convert one that’s 70% paid social, permanently, regardless of how good either is.

3. The sample is self-selected. Most benchmark reports come from a vendor’s customer base. That’s not a random sample of retailers — it’s the retailers who bought that tool, which correlates with size, sophistication and budget. Whether that biases up or down is unknowable, which is the problem.

4. The publisher has an interest. Benchmarks are marketing. A figure showing the industry average at 2.6% while implying their customers reach 4% is a sales argument in the shape of a statistic. This doesn’t make the number false; it means the methodology deserves reading before the figure is quoted.

What to check before quoting one

□  what exactly is the numerator and denominator?
□  what population — how many sites, which sectors, which countries?
□  self-selected, or sampled?
□  what date range, and does it span a peak?
□  mean or median?              ← mean is inflated by a few huge retailers
□  who published it, and what do they sell?
□  is the underlying data available, or just the headline?

If more than two are unanswerable, the figure isn’t usable for comparison — though it may still be usable for the one thing benchmarks are reliably good at, below.

Prefer medians over means wherever offered. Commerce metrics are heavily skewed, so the mean describes a business much larger than typical — Skewed and Heavy-Tailed Distributions, Mean Median and Mode.

What benchmarks are legitimately good for

Narrow uses where the definitional problem doesn’t bite:

  • Shape, not level. Industry seasonal curves, weekday patterns, device mix trends. You’re comparing a pattern rather than a value, and patterns survive definitional differences — Seasonality
  • Order-of-magnitude sanity checks. If published rates cluster at 1–4% and yours reads 22%, you have a measurement bug, not a triumph. This is genuinely valuable and catches real instrumentation faults — the symptom list
  • Direction of travel. “Mobile share of commerce traffic has been rising” is robust to definitional noise in a way “mobile converts at 1.8%” isn’t
  • Standards where the definition is fixed by someone else. Core Web Vitals thresholds are defined identically for everyone, so comparison is meaningful — a rare and useful exception — Core Web Vitals, Field vs Lab Data

The comparison that actually works

Benchmark against yourself.

                    conversion rate     vs
this month              4.2%            last month, same segment
                                        same month last year
                                        the same page pre-redesign
                                        your other brand / market
                                        your best-performing segment

Internal comparison holds every definition constant — same tracking, same population, same denominator — which is precisely what external benchmarking cannot do. A gap between your mobile and desktop conversion, or between your best and worst category page, is a real, actionable, defensible number.

Your own best segment is the most useful benchmark available. If your returning-customer conversion is 9% and new-customer is 2%, the gap is a genuine opportunity measured on identical definitions. No industry report improves on that.

Using them politically

The realistic situation is being handed a benchmark by someone senior. Two responses that work better than rejecting it:

  • Ask what decision it would change. Usually none, which ends it constructively
  • Reframe to an internal comparison. “The industry figure isn’t comparable to our definition, but here’s our trend and the gap between our best and worst segments” — which gives them the reassurance they were seeking, from a defensible number

Never adopt an external benchmark as a target. It imports someone else’s definition and someone else’s population into your objectives, and the first thing that happens is pressure to redefine your metric until it matches — North Star Metric, Vanity Metrics.

Where it interacts

  • Metric Design — the definitional problem that makes most benchmarking invalid, stated in general
  • Tool Discrepancies — the same problem inside one organisation: two tools, two definitions, one argument
  • Communicating Uncertainty — a benchmark quoted without its methodology carries false authority, and stating the caveat is the job
  • Cohort Analysis — the internal comparison that most reliably replaces an external one