Tags: web-dev statistics concept
Percentiles in Performance
Date: 2026-08-16
Performance distributions have a long right tail, so the mean describes nobody. The 75th percentile is the standard because it represents a real user having a bad-but-not-extreme time — and it’s the number that moves when you fix things.
What it is
A percentile is the value below which a given proportion of observations fall. The 75th percentile of LCP is the time by which 75% of page loads had painted their main content.
Percentile-based reporting is the convention in performance because the underlying distribution is heavily right-skewed — see Skewed and Heavy-Tailed Distributions.
Why the mean lies
Ten page loads, LCP in seconds:
1.1 1.2 1.2 1.3 1.4 1.5 1.6 1.8 4.2 11.0
mean 2.63s ← no single load was near this
median 1.45s ← the typical experience
p75 1.75s ← a realistic bad-ish experience
p95 ~9.6s ← the tail, where users leave
The mean of 2.63 is dragged up by one 11-second load and describes nothing that happened. In plain terms: performance data isn’t a bell curve — most loads cluster low and a few are catastrophically slow, so averaging mixes two different populations into one number that represents neither.
Worse, the mean is unstable. Add one more 11-second load and it jumps; the median barely moves. A metric that swings on individual outliers can’t be tracked over time.
Why p75 specifically
Core Web Vitals are assessed at the 75th percentile, which is a deliberate compromise:
- p50 (median) is too forgiving. Half your users are worse than this, and the half that leaves is in the other half
- p75 means three in four visits meet the threshold. Enough of a bar to represent real users on real devices, not just the lucky ones
- p95 / p99 are dominated by genuine outliers — a phone in a lift, a device with 40 tabs open. Chasing them costs a lot and helps few
The consequence worth internalising: passing at p75 still means a quarter of your visits are worse than the threshold. That’s not failure, it’s what the standard accepts.
Percentiles don’t add up
A genuine trap. You cannot combine percentiles the way you combine averages.
p75 of category pages 2.1s
p75 of product pages 2.4s
p75 of the site ??? ← NOT 2.25s
The site’s p75 depends on the whole combined distribution and the traffic mix between templates. Averaging two percentiles gives a number that means nothing.
The same applies to phases. The p75 of LCP is not the sum of the p75s of its four phases — a load with a slow server usually isn’t the same load that had a slow image. Sum the phases within each load, then take the percentile.
Reading a distribution properly
Look at the shape, not one number:
- A wide gap between p50 and p75 means the experience is inconsistent — often two populations, like cached and uncached, or mobile and desktop
- A long tail beyond p95 points at a specific failing segment — one country, one device class, one template
- A bimodal distribution almost always means something binary: cache hit versus miss, consent granted versus denied, one CDN region misconfigured
Segment before optimising. A site-wide p75 of 2.8s might be desktop at 1.6s and mobile at 4.1s, which are two completely different projects.
Practical
- Report p75 as standard, p50 alongside for context, and p95 when hunting a tail
- Never report a mean. If a dashboard shows one, it’s the wrong dashboard
- Segment by device, template, country and connection before drawing conclusions
- Watch the histogram, not just the number. CrUX and most real user monitoring (RUM) tools expose the good / needs-improvement / poor split — that’s a three-bucket histogram and it’s more informative than a single value
- Sample size matters. A p75 on 40 page views is noise. Low-traffic templates need a longer window before their numbers mean anything — Sampling Error