Tags: statistics concept
Probability Distributions
Date: 2026-08-16
The shape of how often each value occurs. It’s not decoration — the shape decides which method is valid, and assuming the wrong one is how confident wrong answers get produced.
What it is
A probability distribution describes how likely each possible value of a random variable is.
NORMAL SKEWED (revenue) BINARY (converted?)
╱▔▔╲ ▌ ▌
╱ ╲ ▌▖ ▌
╱ ╲ ▌▝▖▁▁▁▁▁▁▁▁ ▌ ▌
╱ ╲ ▌ ▔▔▔▔▁▁▁ ▌ ▌
──────────── ────────────────── ────────────
symmetric long right tail two values
The three shapes you’ll actually meet, and each demands different handling.
The ones that matter here
| Distribution | Describes | Where |
|---|---|---|
| Bernoulli | One yes/no trial | Did this visitor convert? |
| Binomial | Count of successes in n trials | How many of 53,000 converted — Binomial and Bernoulli Distributions |
| Normal | Symmetric, bell-shaped | Sample averages, via the CLT |
| Log-normal / heavy-tailed | Multiplicative, long right tail | Order value, session duration, page load — Skewed and Heavy-Tailed Distributions |
| Poisson | Counts of rare events per period | Complaints per week, errors per hour |
Why the shape decides the method
Most standard tests assume normality somewhere — usually in the sampling distribution of the statistic rather than in the raw data. Get that wrong and the intervals and p-values are wrong in ways nothing warns you about.
conversion rate raw data is Bernoulli (0 or 1)
the SAMPLE MEAN goes normal at scale
→ normal-approximation tests are fine
revenue per visitor raw data is heavy-tailed
the sample mean goes normal SLOWLY
→ at modest samples the approximation
is poor, and intervals are too narrow
In plain terms: the maths doesn’t need your data to be bell-shaped. It needs the average of your data to be, and how quickly that happens depends on how skewed the underlying data is. Revenue is the case where it doesn’t happen fast enough.
That’s the practical consequence, and it’s the reason Bootstrapping exists — it makes no distributional assumption at all.
Describing one
Four properties, in the order you’d check them:
- Centre — mean, median. They diverge as skew increases, and the divergence is itself the signal — Mean Median and Mode
- Spread — variance, standard deviation — Variance and Standard Deviation
- Shape — symmetric or skewed, one peak or several
- Tails — how much probability sits far from the centre. Heavy tails are what break averages
A bimodal distribution — two peaks — almost always means two populations mixed together. Cache hit versus miss, mobile versus desktop, consenting versus not. That’s a segmentation finding, not a distribution to model — Segmentation (analysis).
Looking before you calculate
The habit worth building: plot it before you summarise it.
A mean and a standard deviation describe a normal distribution completely and describe a skewed one misleadingly. Any dataset can produce those two numbers; only some are described by them.
select width_bucket(order_value, 0, 500, 50) as bucket,
count(*)
from orders group by 1 order by 1Thirty seconds, and it tells you whether the mean means anything, whether there are two populations, and whether the tail is heavy enough to need Winsorisation and Capping.
Where it feeds
- Sample Size Calculation — the variance term comes from the distribution
- Percentiles in Performance — why p75 rather than the mean, for anything skewed
- Metric Sensitivity — the coefficient of variation, which is a distribution property
- Outliers and Robust Statistics — what to do about the tail