Tags: statistics concept
Percentiles and Quantiles
Date: 2026-08-16
Reading a distribution by position rather than by average. It’s the honest way to describe anything skewed — and the one property that trips people is that percentiles can’t be averaged or added.
What it is
A percentile is the value below which a given proportion of observations fall. The 75th percentile of order value is the amount 75% of orders came in under.
Quantile is the general term; percentiles are quantiles in hundredths. Quartiles are quarters, deciles tenths.
sorted order values
£18 £24 £28 £31 £35 £38 £42 £48 £55 £2,400
│ │
p50 p90
£36.50 £790
mean £272.30 — sits at the 90th percentile.
Nine orders in ten were below the "average".
That last line is the whole argument for percentiles.
Where each is useful
| Reads as | |
|---|---|
| p50 (median) | Typical |
| p75 | A realistically bad case. The Core Web Vitals standard |
| p90 / p95 | The tail — where users leave, where systems fail |
| p99 | Genuine outliers. Expensive to chase, affects few |
| IQR (p75 − p25) | Spread, robust to outliers, unlike standard deviation |
p75 is the convention in performance because it’s a compromise: p50 is too forgiving (half your users are worse), p95 is dominated by genuine anomalies. See Percentiles in Performance.
Percentiles don’t add or average
The property that causes real errors.
p75 of category pages 2.1s
p75 of product pages 2.4s
p75 of the site ??? ← NOT 2.25s
The site’s p75 depends on the combined distribution and the traffic mix between templates. Averaging two percentiles produces a number that means nothing.
Same for components. The p75 of a total is not the sum of the p75s of its parts, because the slowest server response usually isn’t on the same request as the slowest image. Sum within each observation, then take the percentile of the totals.
In plain terms: a percentile is a position in a sorted list. You can’t add positions.
Computing them
Several definitions exist and they differ on small samples — nearest rank, linear interpolation, and variants. Tools disagree in the last decimal, which is a real source of “why don’t these match” on small datasets and irrelevant on large ones.
select
percentile_cont(0.5) within group (order by value) as p50,
percentile_cont(0.75) within group (order by value) as p75,
percentile_cont(0.95) within group (order by value) as p95
from orderspercentile_cont interpolates; percentile_disc returns an actual observed value. Use disc when the answer must be a real data point.
Confidence intervals on percentiles
The CLT is about means, so the standard formulas don’t apply to a median or a p75. Getting an interval around a percentile means Bootstrapping — resample and observe how much the percentile moves.
Worth knowing because it’s the usual reason a dashboard shows a p75 with no uncertainty at all: it’s harder to compute, so it gets omitted.
Practical
- Report p50 and p75 together for anything skewed. The gap is the shape
- Never report a mean for performance data. If a dashboard shows one, it’s the wrong dashboard
- Segment before percentiling. A site-wide p75 across mobile and desktop describes neither
- Watch sample size. A p95 on 40 observations is two data points — Sampling Error
- Percentiles of a percentile are meaningless. “Average p75 across the month” is not a thing; compute the p75 over the month’s data