Tags: statistics concept
Random Variables
Date: 2026-08-17
A random variable is a number whose value comes out of a process you can describe but not predict — one visitor’s spend, whether a session converts, how long a page takes to load. The idea matters because it’s what lets you talk about a quantity’s behaviour before you’ve observed it.
The shift in thinking
NOT A RANDOM VARIABLE RANDOM VARIABLE
"3.2% converted last month" "X = whether the NEXT visitor converts"
a fact. already happened. X isn't 0 and isn't 1 yet.
one number. it's a description of a process:
P(X = 1) = 0.032
P(X = 0) = 0.968
The random variable is the process, not the outcome. Once you observe it you have a realisation — a plain number. The distinction sounds pedantic and is what makes the rest of statistics expressible: you can compute the expected behaviour of X before running the experiment, which is exactly what a sample size calculation does.
Convention: capital letters for the variable (X), lowercase for an observed value (x).
Discrete and continuous
| Discrete | Continuous | |
|---|---|---|
| Values | Countable — 0, 1, 2… | Any value in a range |
| Examples here | Converted or not; items in basket; orders per day | Order value; page load time; time on site |
| Described by | Probability of each value | Probability of falling in a range |
P(X = exact value) | Meaningful | Zero — ask for a range instead |
The last row trips people up: for a continuous variable, the probability of load time being exactly 1.4000… seconds is zero. Only P(1.3 < X < 1.5) is meaningful.
The two numbers that describe one
Expectation E[X] — the long-run average, weighting each value by how often it occurs.
E[X] = Σ (value × probability of that value)
Variance Var(X) — the average squared distance from the expectation.
Var(X) = E[(X − E[X])²]
Worked, for the most important case in this vault — a Bernoulli variable, meaning a single yes/no outcome with probability p:
X = 1 if a visitor converts, 0 if not. p = 0.03
E[X] = (1 × 0.03) + (0 × 0.97)
= 0.03 ← the conversion rate IS the expectation
Var(X) = E[X²] − (E[X])²
= [(1² × 0.03) + (0² × 0.97)] − 0.03²
= 0.03 − 0.0009
= 0.0291
SD(X) = √0.0291 = 0.171
In plain terms: the average outcome per visitor is 0.03 conversions, and the typical distance of any single visitor from that average is 0.171 — which is enormous relative to 0.03, because every individual outcome is either 0 or 1 and neither is anywhere near 0.03. That ratio is why single observations tell you nothing and why conversion tests need tens of thousands of visitors.
For a Bernoulli variable this simplifies to a formula worth memorising, since it appears in every sample size calculation:
Var(X) = p(1 − p) = 0.03 × 0.97 = 0.0291 ✓ matches above
— Binomial and Bernoulli Distributions.
Adding them up, which is where it pays off
Two rules do most of the work in practice. For random variables X and Y:
E[X + Y] = E[X] + E[Y] always true, even if X and Y are related
Var(X + Y) = Var(X) + Var(Y) only if X and Y are INDEPENDENT
The condition on the second is where real analyses break. Sessions from the same user are not independent — someone who converts once is more likely to convert again — so treating 100,000 sessions as 100,000 independent draws understates the variance and produces confidence intervals that are too narrow.
This is exactly why the randomisation unit matters: randomise by user, analyse by user, or the independence assumption fails and every interval you compute is wrong in the optimistic direction.
Worked, for a total across a sample:
1,000 independent visitors, each Bernoulli with p = 0.03
S = total conversions
E[S] = 1,000 × 0.03 = 30
Var(S) = 1,000 × 0.0291 = 29.1
SD(S) = √29.1 = 5.4
so a typical observed count is 30 ± 5.4 — meaning 25 or 36
conversions from 1,000 visitors is entirely ordinary, and
reading a difference that size as a change is reading noise
Where this leads
- A sample mean is itself a random variable, with its own distribution, its own expectation and its own variance — and that variance is what Standard Error measures
- That distribution goes normal as the sample grows, whatever the underlying variable looks like — The Central Limit Theorem
- Which is why formulas assuming normality work on conversion rates, even though a single conversion is as un-normal as a variable can be
Where it interacts
- Probability Distributions — the catalogue of shapes a random variable can take
- Variance and Standard Deviation — the spread measure defined here, used everywhere
- Ratio Metrics — a ratio of two random variables has a variance that isn’t derivable from its parts, which is the practical consequence of the addition rules above
- Metric Sensitivity — driven entirely by the variance-to-mean ratio computed here