Tags: statistics concept

Random Variables

Date: 2026-08-17


A random variable is a number whose value comes out of a process you can describe but not predict — one visitor’s spend, whether a session converts, how long a page takes to load. The idea matters because it’s what lets you talk about a quantity’s behaviour before you’ve observed it.


The shift in thinking

NOT A RANDOM VARIABLE                RANDOM VARIABLE

"3.2% converted last month"          "X = whether the NEXT visitor converts"

a fact. already happened.            X isn't 0 and isn't 1 yet.
one number.                          it's a description of a process:
                                       P(X = 1) = 0.032
                                       P(X = 0) = 0.968

The random variable is the process, not the outcome. Once you observe it you have a realisation — a plain number. The distinction sounds pedantic and is what makes the rest of statistics expressible: you can compute the expected behaviour of X before running the experiment, which is exactly what a sample size calculation does.

Convention: capital letters for the variable (X), lowercase for an observed value (x).

Discrete and continuous

DiscreteContinuous
ValuesCountable — 0, 1, 2…Any value in a range
Examples hereConverted or not; items in basket; orders per dayOrder value; page load time; time on site
Described byProbability of each valueProbability of falling in a range
P(X = exact value)MeaningfulZero — ask for a range instead

The last row trips people up: for a continuous variable, the probability of load time being exactly 1.4000… seconds is zero. Only P(1.3 < X < 1.5) is meaningful.

The two numbers that describe one

Expectation E[X] — the long-run average, weighting each value by how often it occurs.

E[X]  =  Σ  (value × probability of that value)

Variance Var(X) — the average squared distance from the expectation.

Var(X)  =  E[(X − E[X])²]

Worked, for the most important case in this vault — a Bernoulli variable, meaning a single yes/no outcome with probability p:

X = 1 if a visitor converts, 0 if not.   p = 0.03

E[X]    =  (1 × 0.03) + (0 × 0.97)
        =  0.03                          ← the conversion rate IS the expectation

Var(X)  =  E[X²] − (E[X])²
        =  [(1² × 0.03) + (0² × 0.97)] − 0.03²
        =  0.03 − 0.0009
        =  0.0291

SD(X)   =  √0.0291  =  0.171

In plain terms: the average outcome per visitor is 0.03 conversions, and the typical distance of any single visitor from that average is 0.171 — which is enormous relative to 0.03, because every individual outcome is either 0 or 1 and neither is anywhere near 0.03. That ratio is why single observations tell you nothing and why conversion tests need tens of thousands of visitors.

For a Bernoulli variable this simplifies to a formula worth memorising, since it appears in every sample size calculation:

Var(X) = p(1 − p)  =  0.03 × 0.97  =  0.0291   ✓ matches above

— Binomial and Bernoulli Distributions.

Adding them up, which is where it pays off

Two rules do most of the work in practice. For random variables X and Y:

E[X + Y]  =  E[X] + E[Y]           always true, even if X and Y are related

Var(X + Y) =  Var(X) + Var(Y)      only if X and Y are INDEPENDENT

The condition on the second is where real analyses break. Sessions from the same user are not independent — someone who converts once is more likely to convert again — so treating 100,000 sessions as 100,000 independent draws understates the variance and produces confidence intervals that are too narrow.

This is exactly why the randomisation unit matters: randomise by user, analyse by user, or the independence assumption fails and every interval you compute is wrong in the optimistic direction.

Worked, for a total across a sample:

1,000 independent visitors, each Bernoulli with p = 0.03
S = total conversions

E[S]    =  1,000 × 0.03      =  30
Var(S)  =  1,000 × 0.0291    =  29.1
SD(S)   =  √29.1             =  5.4

so a typical observed count is 30 ± 5.4 — meaning 25 or 36
conversions from 1,000 visitors is entirely ordinary, and
reading a difference that size as a change is reading noise

Where this leads

  • A sample mean is itself a random variable, with its own distribution, its own expectation and its own variance — and that variance is what Standard Error measures
  • That distribution goes normal as the sample grows, whatever the underlying variable looks like — The Central Limit Theorem
  • Which is why formulas assuming normality work on conversion rates, even though a single conversion is as un-normal as a variable can be

Where it interacts