Tags: statistics concept

Probability Distributions

Date: 2026-08-16


The shape of how often each value occurs. It’s not decoration — the shape decides which method is valid, and assuming the wrong one is how confident wrong answers get produced.


What it is

A probability distribution describes how likely each possible value of a random variable is.

NORMAL                  SKEWED (revenue)         BINARY (converted?)

    ╱▔▔╲                 ▌                        ▌
   ╱    ╲                ▌▖                       ▌
  ╱      ╲               ▌▝▖▁▁▁▁▁▁▁▁              ▌         ▌
 ╱        ╲              ▌         ▔▔▔▔▁▁▁        ▌         ▌
────────────            ──────────────────       ────────────
  symmetric              long right tail          two values

The three shapes you’ll actually meet, and each demands different handling.

The ones that matter here

DistributionDescribesWhere
BernoulliOne yes/no trialDid this visitor convert?
BinomialCount of successes in n trialsHow many of 53,000 converted — Binomial and Bernoulli Distributions
NormalSymmetric, bell-shapedSample averages, via the CLT
Log-normal / heavy-tailedMultiplicative, long right tailOrder value, session duration, page load — Skewed and Heavy-Tailed Distributions
PoissonCounts of rare events per periodComplaints per week, errors per hour

Why the shape decides the method

Most standard tests assume normality somewhere — usually in the sampling distribution of the statistic rather than in the raw data. Get that wrong and the intervals and p-values are wrong in ways nothing warns you about.

conversion rate      raw data is Bernoulli (0 or 1)
                     the SAMPLE MEAN goes normal at scale
                     → normal-approximation tests are fine

revenue per visitor  raw data is heavy-tailed
                     the sample mean goes normal SLOWLY
                     → at modest samples the approximation
                       is poor, and intervals are too narrow

In plain terms: the maths doesn’t need your data to be bell-shaped. It needs the average of your data to be, and how quickly that happens depends on how skewed the underlying data is. Revenue is the case where it doesn’t happen fast enough.

That’s the practical consequence, and it’s the reason Bootstrapping exists — it makes no distributional assumption at all.

Describing one

Four properties, in the order you’d check them:

  • Centre — mean, median. They diverge as skew increases, and the divergence is itself the signal — Mean Median and Mode
  • Spread — variance, standard deviation — Variance and Standard Deviation
  • Shape — symmetric or skewed, one peak or several
  • Tails — how much probability sits far from the centre. Heavy tails are what break averages

A bimodal distribution — two peaks — almost always means two populations mixed together. Cache hit versus miss, mobile versus desktop, consenting versus not. That’s a segmentation finding, not a distribution to model — Segmentation (analysis).

Looking before you calculate

The habit worth building: plot it before you summarise it.

A mean and a standard deviation describe a normal distribution completely and describe a skewed one misleadingly. Any dataset can produce those two numbers; only some are described by them.

select width_bucket(order_value, 0, 500, 50) as bucket,
       count(*)
from orders group by 1 order by 1

Thirty seconds, and it tells you whether the mean means anything, whether there are two populations, and whether the tail is heavy enough to need Winsorisation and Capping.

Where it feeds