Tags: web-dev statistics concept

Percentiles in Performance

Date: 2026-08-16


Performance distributions have a long right tail, so the mean describes nobody. The 75th percentile is the standard because it represents a real user having a bad-but-not-extreme time — and it’s the number that moves when you fix things.


What it is

A percentile is the value below which a given proportion of observations fall. The 75th percentile of LCP is the time by which 75% of page loads had painted their main content.

Percentile-based reporting is the convention in performance because the underlying distribution is heavily right-skewed — see Skewed and Heavy-Tailed Distributions.

Why the mean lies

Ten page loads, LCP in seconds:

1.1  1.2  1.2  1.3  1.4  1.5  1.6  1.8  4.2  11.0

mean    2.63s      ← no single load was near this
median  1.45s      ← the typical experience
p75     1.75s      ← a realistic bad-ish experience
p95     ~9.6s      ← the tail, where users leave

The mean of 2.63 is dragged up by one 11-second load and describes nothing that happened. In plain terms: performance data isn’t a bell curve — most loads cluster low and a few are catastrophically slow, so averaging mixes two different populations into one number that represents neither.

Worse, the mean is unstable. Add one more 11-second load and it jumps; the median barely moves. A metric that swings on individual outliers can’t be tracked over time.

Why p75 specifically

Core Web Vitals are assessed at the 75th percentile, which is a deliberate compromise:

  • p50 (median) is too forgiving. Half your users are worse than this, and the half that leaves is in the other half
  • p75 means three in four visits meet the threshold. Enough of a bar to represent real users on real devices, not just the lucky ones
  • p95 / p99 are dominated by genuine outliers — a phone in a lift, a device with 40 tabs open. Chasing them costs a lot and helps few

The consequence worth internalising: passing at p75 still means a quarter of your visits are worse than the threshold. That’s not failure, it’s what the standard accepts.

Percentiles don’t add up

A genuine trap. You cannot combine percentiles the way you combine averages.

p75 of category pages   2.1s
p75 of product pages    2.4s
p75 of the site         ???     ← NOT 2.25s

The site’s p75 depends on the whole combined distribution and the traffic mix between templates. Averaging two percentiles gives a number that means nothing.

The same applies to phases. The p75 of LCP is not the sum of the p75s of its four phases — a load with a slow server usually isn’t the same load that had a slow image. Sum the phases within each load, then take the percentile.

Reading a distribution properly

Look at the shape, not one number:

  • A wide gap between p50 and p75 means the experience is inconsistent — often two populations, like cached and uncached, or mobile and desktop
  • A long tail beyond p95 points at a specific failing segment — one country, one device class, one template
  • A bimodal distribution almost always means something binary: cache hit versus miss, consent granted versus denied, one CDN region misconfigured

Segment before optimising. A site-wide p75 of 2.8s might be desktop at 1.6s and mobile at 4.1s, which are two completely different projects.

Practical

  • Report p75 as standard, p50 alongside for context, and p95 when hunting a tail
  • Never report a mean. If a dashboard shows one, it’s the wrong dashboard
  • Segment by device, template, country and connection before drawing conclusions
  • Watch the histogram, not just the number. CrUX and most real user monitoring (RUM) tools expose the good / needs-improvement / poor split — that’s a three-bucket histogram and it’s more informative than a single value
  • Sample size matters. A p75 on 40 page views is noise. Low-traffic templates need a longer window before their numbers mean anything — Sampling Error