Tags: statistics concept

Linear Regression

Date: 2026-08-17


Fitting a straight line through data so the relationship has a slope in real units. Where Correlation says “these move together”, regression says “each extra second of load time costs 0.44 percentage points” — which is the form a decision can be made from.


Linear regression fits the straight line that best predicts one variable from another, giving the relationship as a slope in the data’s own units.

The model

y  =  a  +  b·x  +  error

a  intercept — predicted y when x = 0
b  slope     — change in y per one-unit change in x     ← the useful number

Fitted by least squares: choose a and b to minimise the sum of squared vertical distances from the points to the line.

Worked

Same five page templates as in Correlation — load time against conversion rate:

x (load, s)   y (conv %)    dx      dy      dx·dy    dx²
   1.0          4.0        −2.0    +0.8     −1.60    4.0
   2.0          3.6        −1.0    +0.4     −0.40    1.0
   3.0          3.4         0.0    +0.2      0.00    0.0
   4.0          2.8        +1.0    −0.4     −0.40    1.0
   5.0          2.2        +2.0    −1.0     −2.00    4.0
                                          ───────  ─────
x̄ = 3.0     ȳ = 3.2                        −4.40   10.0

b  =  Σ(dx·dy) / Σdx²   =  −4.40 / 10.0   =  −0.44
a  =  ȳ − b·x̄           =  3.2 − (−0.44 × 3.0)  =  3.2 + 1.32  =  4.52

           ŷ  =  4.52  −  0.44x

Reading it: each additional second of load time is associated with a 0.44 percentage-point fall in conversion rate. At a 3% baseline that’s roughly a 15% relative loss per second — a number you can put against the cost of an optimisation project — Performance and Conversion.

Checking the fit:

x     actual y    predicted ŷ    residual
1.0     4.0         4.08          −0.08
2.0     3.6         3.64          −0.04
3.0     3.4         3.20          +0.20
4.0     2.8         2.76          +0.04
5.0     2.2         2.32          −0.12

Residuals are small and don’t trend — the line describes this data well. r² = 0.968, so ~97% of the variation in conversion across these templates is accounted for by load time.

In plain terms: on these five templates, knowing the load time lets you guess the conversion rate almost exactly. That says the line fits; it says nothing about whether it would hold on a sixth template, or whether load time is the cause.

What the intercept usually isn’t

a = 4.52 predicts conversion at zero load time — physically impossible, and outside the observed range of 1–5 seconds.

Never extrapolate beyond your data. The relationship is fitted where you have observations; nothing licenses it outside. A model saying conversion hits 4.5% at instantaneous load is arithmetic, not a finding.

The assumptions, and which actually bite

AssumptionWhat breaks if violatedHow much it matters
LinearityThe whole model is wrongCritically. Plot first, always
IndependenceStandard errors far too smallCritically — repeated sessions per user break this
Constant variance (homoscedasticity)Intervals wrong, coefficients fineModerately
Normal residualsSmall-sample intervals wrongBarely, at analytics n
No extreme outliersOne point moves the lineSubstantially — see below

Independence is the one that quietly ruins commercial analyses. Fitting a regression across 200,000 sessions from 80,000 users treats correlated observations as independent, understating the standard errors — so everything looks more significant than it is. Aggregate to the user, or use a model that accounts for clustering — Random Variables, Randomisation Unit.

Plot the residuals against the fitted values. A curve means non-linearity; a fan shape means non-constant variance. This one chart catches most problems and takes a second.

Outliers move the line

Least squares minimises squared errors, so a point twice as far away has four times the influence.

add one page: load 12.0s, conversion 3.9%
(a heavy but well-cached page that converts fine)

original    ŷ = 4.52 − 0.44x      r² = 0.968
with it     ŷ = 3.27 + 0.01x      r² = 0.004

the slope vanished — it even flipped sign — and the model stopped
describing anything, on the strength of one observation

Check influence before believing a slope. Refit without each extreme point; if the conclusion changes, say so rather than picking the version you prefer — Outliers and Robust Statistics.

What it does not establish

A slope is not a causal effect. The −0.44 above is consistent with load time hurting conversion, and equally with heavy pages being the ones that are complex, image-rich or on the wrong templates — all of which independently affect conversion.

Regression on observational data adjusts for what you put in it and nothing else. The causal claim needs a design, not a model — Correlation and Causation, Confounding Variables, Randomised Controlled Trials.

Where it’s genuinely used here

  • Quantifying a relationship in decision units — the pounds-per-second figure above
  • Variance Reduction — CUPED (controlled experiment using pre-experiment data) is regression on a pre-period covariate, and it’s the highest-value application of the technique in experimentation
  • Baselines and forecasting — trend estimation for Seasonality decomposition
  • Marketing Mix Modelling — a regression with adstock and saturation transformations bolted on

Where it interacts

  • Correlation — the same relationship without units; r is the standardised version of b
  • Multiple Regression — more than one predictor, and the illusion of control that creates
  • Logistic Regression — the version for binary outcomes, which is most outcomes in this vault
  • Regression to the Mean — different concept, confusingly similar name, and the source of the word “regression” here