Tags: statistics concept

Correlation

Date: 2026-08-17


A single number summarising how two variables move together. It measures linear co-movement only, it’s badly distorted by outliers, and squaring it gives the more honest figure — an r of 0.7 sounds strong and explains under half the variation.


Pearson’s r, worked

The correlation coefficient, written r, is a number from −1 to +1 measuring how closely two variables track a straight line together. Pearson’s is the standard one — computed from the actual values — as opposed to Spearman’s, which uses ranks instead and appears below.

 r = −1        perfect inverse: every rise in x is a proportional fall in y
 r =  0        no LINEAR relationship — which is not the same as no relationship
 r = +1        perfect direct

sign      the direction
magnitude how tightly the points sit on the line, NOT how steep it is

The magnitude says nothing about the slope. A relationship costing 0.01pp per second and one costing 4pp per second both give r = −1 if they’re equally consistent. For the size of the effect you need a regression — Linear Regression.

        r  =  Σ(dx · dy)  /  √( Σdx² · Σdy² )

where dx = x − x̄  and  dy = y − ȳ

Page load time against conversion rate, across five page templates:

x (load, s)   y (conv %)    dx      dy      dx·dy    dx²    dy²
   1.0          4.0        −2.0    +0.8     −1.60    4.0   0.64
   2.0          3.6        −1.0    +0.4     −0.40    1.0   0.16
   3.0          3.4         0.0    +0.2      0.00    0.0   0.04
   4.0          2.8        +1.0    −0.4     −0.40    1.0   0.16
   5.0          2.2        +2.0    −1.0     −2.00    4.0   1.00
                                          ───────  ─────  ─────
x̄ = 3.0     ȳ = 3.2                        −4.40   10.0   2.00

r  =  −4.40 / √(10.0 × 2.00)
   =  −4.40 / √20
   =  −4.40 / 4.472
   =  −0.984

In plain terms: as load time rises, conversion falls, and it does so almost perfectly consistently across these five templates. The sign says the direction; the magnitude says how tightly the points sit on a straight line.

r² is the number to quote

r²  =  (−0.984)²  =  0.968   →  96.8% of the variation in conversion
                                 is accounted for by load time, in this data

Squaring is what stops r sounding better than it is:

r        r²      variation explained
0.9     0.81            81%
0.7     0.49            49%     ← "strong correlation", under half
0.5     0.25            25%
0.3     0.09             9%     ← "moderate correlation", essentially nothing
0.1     0.01             1%

An r of 0.3 is routinely described as a moderate relationship and explains 9% of the variation. Quoting r² is the single cheapest correction to overclaiming in analytics reporting.

What r cannot see

Non-linearity. r measures straight-line association only.

y = x² over −5 to +5

  x    −5   −3   −1    1    3    5
  y     25    9    1    1    9   25

r = 0    ← perfect deterministic relationship, zero correlation

A relationship where more is better up to a point and then worse — discount depth, email frequency, number of form fields — will show near-zero r while being entirely real and entirely actionable. Always plot before trusting an r.

Outliers. A single extreme point can create or destroy a correlation.

9 points with r = 0.02, plus one extreme point in the corner
  → r jumps to 0.71

remove that one point  →  r returns to 0.02

With heavy-tailed data — order value, session duration, revenue per customer — this is the normal case rather than the exception. Use Spearman’s rank correlation (Pearson’s r computed on ranks) as a robustness check: if Pearson and Spearman disagree sharply, outliers or non-linearity are driving your number — Outliers and Robust Statistics, Skewed and Heavy-Tailed Distributions.

Range restriction. Correlations computed on a narrow slice understate the relationship. If every page on your site loads between 1.8s and 2.2s, load time will correlate weakly with conversion — not because it doesn’t matter, but because you haven’t observed enough variation to see it.

Aggregation. Correlations computed on group averages are typically far stronger than the same relationship at individual level — the ecological fallacy. A 0.9 correlation across ten country-level averages tells you little about individual users.

Statistical significance of r

A correlation can be significant and negligible at the same time, and at analytics volumes this happens constantly.

n = 100,000 sessions
r = 0.01     p < 0.001     ← "highly significant"
             r² = 0.0001   ← explains one hundredth of one percent

In plain terms: with enough data, essentially every correlation is statistically distinguishable from zero. Significance tells you the relationship isn’t exactly nil; r² tells you whether it’s worth anything. Report both, and let r² do the deciding — Practical vs Statistical Significance.

Where it’s genuinely useful

  • Screening for candidate relationships before modelling — Linear Regression
  • Validating a surrogate metric. Does add-to-cart rate correlate with orders across your archive of past tests? That’s the check that licenses using it as a proxy — Metric Selection for Tests
  • Detecting redundancy. Two metrics correlating at 0.95 are one metric, and putting both in a model causes problems — Multiple Regression
  • Variance Reduction. CUPED — controlled experiment using pre-experiment data — has a benefit driven by ρ² — a pre-period correlation of 0.5 cuts required sample by 25%, and this is one place where a modest r has direct cash value

Where it interacts