Tags: statistics concept
Linear Regression
Date: 2026-08-17
Fitting a straight line through data so the relationship has a slope in real units. Where Correlation says “these move together”, regression says “each extra second of load time costs 0.44 percentage points” — which is the form a decision can be made from.
Linear regression fits the straight line that best predicts one variable from another, giving the relationship as a slope in the data’s own units.
The model
y = a + b·x + error
a intercept — predicted y when x = 0
b slope — change in y per one-unit change in x ← the useful number
Fitted by least squares: choose a and b to minimise the sum of squared vertical distances from the points to the line.
Worked
Same five page templates as in Correlation — load time against conversion rate:
x (load, s) y (conv %) dx dy dx·dy dx²
1.0 4.0 −2.0 +0.8 −1.60 4.0
2.0 3.6 −1.0 +0.4 −0.40 1.0
3.0 3.4 0.0 +0.2 0.00 0.0
4.0 2.8 +1.0 −0.4 −0.40 1.0
5.0 2.2 +2.0 −1.0 −2.00 4.0
─────── ─────
x̄ = 3.0 ȳ = 3.2 −4.40 10.0
b = Σ(dx·dy) / Σdx² = −4.40 / 10.0 = −0.44
a = ȳ − b·x̄ = 3.2 − (−0.44 × 3.0) = 3.2 + 1.32 = 4.52
ŷ = 4.52 − 0.44x
Reading it: each additional second of load time is associated with a 0.44 percentage-point fall in conversion rate. At a 3% baseline that’s roughly a 15% relative loss per second — a number you can put against the cost of an optimisation project — Performance and Conversion.
Checking the fit:
x actual y predicted ŷ residual
1.0 4.0 4.08 −0.08
2.0 3.6 3.64 −0.04
3.0 3.4 3.20 +0.20
4.0 2.8 2.76 +0.04
5.0 2.2 2.32 −0.12
Residuals are small and don’t trend — the line describes this data well. r² = 0.968, so ~97% of the variation in conversion across these templates is accounted for by load time.
In plain terms: on these five templates, knowing the load time lets you guess the conversion rate almost exactly. That says the line fits; it says nothing about whether it would hold on a sixth template, or whether load time is the cause.
What the intercept usually isn’t
a = 4.52 predicts conversion at zero load time — physically impossible, and outside the observed range of 1–5 seconds.
Never extrapolate beyond your data. The relationship is fitted where you have observations; nothing licenses it outside. A model saying conversion hits 4.5% at instantaneous load is arithmetic, not a finding.
The assumptions, and which actually bite
| Assumption | What breaks if violated | How much it matters |
|---|---|---|
| Linearity | The whole model is wrong | Critically. Plot first, always |
| Independence | Standard errors far too small | Critically — repeated sessions per user break this |
| Constant variance (homoscedasticity) | Intervals wrong, coefficients fine | Moderately |
| Normal residuals | Small-sample intervals wrong | Barely, at analytics n |
| No extreme outliers | One point moves the line | Substantially — see below |
Independence is the one that quietly ruins commercial analyses. Fitting a regression across 200,000 sessions from 80,000 users treats correlated observations as independent, understating the standard errors — so everything looks more significant than it is. Aggregate to the user, or use a model that accounts for clustering — Random Variables, Randomisation Unit.
Plot the residuals against the fitted values. A curve means non-linearity; a fan shape means non-constant variance. This one chart catches most problems and takes a second.
Outliers move the line
Least squares minimises squared errors, so a point twice as far away has four times the influence.
add one page: load 12.0s, conversion 3.9%
(a heavy but well-cached page that converts fine)
original ŷ = 4.52 − 0.44x r² = 0.968
with it ŷ = 3.27 + 0.01x r² = 0.004
the slope vanished — it even flipped sign — and the model stopped
describing anything, on the strength of one observation
Check influence before believing a slope. Refit without each extreme point; if the conclusion changes, say so rather than picking the version you prefer — Outliers and Robust Statistics.
What it does not establish
A slope is not a causal effect. The −0.44 above is consistent with load time hurting conversion, and equally with heavy pages being the ones that are complex, image-rich or on the wrong templates — all of which independently affect conversion.
Regression on observational data adjusts for what you put in it and nothing else. The causal claim needs a design, not a model — Correlation and Causation, Confounding Variables, Randomised Controlled Trials.
Where it’s genuinely used here
- Quantifying a relationship in decision units — the pounds-per-second figure above
- Variance Reduction — CUPED (controlled experiment using pre-experiment data) is regression on a pre-period covariate, and it’s the highest-value application of the technique in experimentation
- Baselines and forecasting — trend estimation for Seasonality decomposition
- Marketing Mix Modelling — a regression with adstock and saturation transformations bolted on
Where it interacts
- Correlation — the same relationship without units; r is the standardised version of b
- Multiple Regression — more than one predictor, and the illusion of control that creates
- Logistic Regression — the version for binary outcomes, which is most outcomes in this vault
- Regression to the Mean — different concept, confusingly similar name, and the source of the word “regression” here