Tags: statistics concept

Percentiles and Quantiles

Date: 2026-08-16


Reading a distribution by position rather than by average. It’s the honest way to describe anything skewed — and the one property that trips people is that percentiles can’t be averaged or added.


What it is

A percentile is the value below which a given proportion of observations fall. The 75th percentile of order value is the amount 75% of orders came in under.

Quantile is the general term; percentiles are quantiles in hundredths. Quartiles are quarters, deciles tenths.

sorted order values

£18  £24  £28  £31  £35  £38  £42  £48  £55  £2,400
                      │                        │
                    p50                       p90
                   £36.50                    £790

mean £272.30 — sits at the 90th percentile.
Nine orders in ten were below the "average".

That last line is the whole argument for percentiles.

Where each is useful

Reads as
p50 (median)Typical
p75A realistically bad case. The Core Web Vitals standard
p90 / p95The tail — where users leave, where systems fail
p99Genuine outliers. Expensive to chase, affects few
IQR (p75 − p25)Spread, robust to outliers, unlike standard deviation

p75 is the convention in performance because it’s a compromise: p50 is too forgiving (half your users are worse), p95 is dominated by genuine anomalies. See Percentiles in Performance.

Percentiles don’t add or average

The property that causes real errors.

p75 of category pages    2.1s
p75 of product pages     2.4s
p75 of the site          ???   ← NOT 2.25s

The site’s p75 depends on the combined distribution and the traffic mix between templates. Averaging two percentiles produces a number that means nothing.

Same for components. The p75 of a total is not the sum of the p75s of its parts, because the slowest server response usually isn’t on the same request as the slowest image. Sum within each observation, then take the percentile of the totals.

In plain terms: a percentile is a position in a sorted list. You can’t add positions.

Computing them

Several definitions exist and they differ on small samples — nearest rank, linear interpolation, and variants. Tools disagree in the last decimal, which is a real source of “why don’t these match” on small datasets and irrelevant on large ones.

select
  percentile_cont(0.5)  within group (order by value) as p50,
  percentile_cont(0.75) within group (order by value) as p75,
  percentile_cont(0.95) within group (order by value) as p95
from orders

percentile_cont interpolates; percentile_disc returns an actual observed value. Use disc when the answer must be a real data point.

Confidence intervals on percentiles

The CLT is about means, so the standard formulas don’t apply to a median or a p75. Getting an interval around a percentile means Bootstrapping — resample and observe how much the percentile moves.

Worth knowing because it’s the usual reason a dashboard shows a p75 with no uncertainty at all: it’s harder to compute, so it gets omitted.

Practical

  • Report p50 and p75 together for anything skewed. The gap is the shape
  • Never report a mean for performance data. If a dashboard shows one, it’s the wrong dashboard
  • Segment before percentiling. A site-wide p75 across mobile and desktop describes neither
  • Watch sample size. A p95 on 40 observations is two data points — Sampling Error
  • Percentiles of a percentile are meaningless. “Average p75 across the month” is not a thing; compute the p75 over the month’s data