Tags: statistics concept
Multiple Regression
Date: 2026-08-17
Fitting several predictors at once, so each coefficient is the effect of that variable “holding the others constant”. The phrase is what makes it dangerous — it sounds like experimental control and isn’t, because it only holds constant the things you measured and put in the model.
Multiple regression is linear regression with more than one predictor, estimating each predictor’s slope while the others are held fixed in the model.
The model
y = a + b₁x₁ + b₂x₂ + b₃x₃ + error
Each b is the change in y per unit of that x, with the other predictors held fixed.
conversion rate = 4.9 − 0.31·(load seconds)
+ 1.42·(returning customer, 0/1)
− 0.08·(form fields)
b₁ = −0.31 each extra second costs 0.31pp, among customers
of the same returning-status and form length
b₂ = +1.42 returning customers convert 1.42pp higher, at the
same load time and form length
b₃ = −0.08 each extra form field costs 0.08pp, all else equal
Note that b₁ fell from −0.44 in the simple regression to −0.31 here. The difference is what the other variables were absorbing — part of what looked like a load-time effect was actually returning customers being served faster pages.
”Controlling for” is not control
The central point of the note.
EXPERIMENTAL CONTROL STATISTICAL CONTROL
randomise add a column to a regression
balances EVERY variable — adjusts for the variables you
measured, unmeasured, and measured, specified correctly,
never imagined and modelled with the right
functional form
— Why Randomisation Works everything else remains confounded
Randomisation achieves the left-hand column’s guarantee without needing the list at all — Why Randomisation Works.
In plain terms: a regression controlling for device, channel and returning-status has not controlled for purchase intent, urgency, income, whether the customer had already decided, or anything else you don’t have a column for. Those are usually the variables that matter most, and they’re unmeasured precisely because they’re hard to measure — Confounding Variables.
The language does real damage here. “We controlled for X” is heard as “X has been eliminated as an explanation”, and what it means is “X is one of the several dozen possible explanations we were able to adjust for”.
Three ways adding variables makes it worse
1. Multicollinearity. Predictors that move together can’t be separated.
model includes both "page weight (KB)" and "number of images"
correlation between them: 0.94
fitted coefficients
page weight −0.002 (SE 0.004) ← "not significant"
images −0.140 (SE 0.210) ← "not significant"
drop one, and the other becomes strongly significant.
the model cannot attribute the effect between two variables
that always move together. it isn't confused — the DATA
contains no information to separate them
Symptoms: large standard errors, coefficients that flip sign when a variable is added or removed, a model with high overall r² and no individually significant predictors.
2. Colliders. Adding a variable caused by both the predictor and the outcome creates a spurious association that wasn’t there.
regressing conversion on treatment, "controlling for" whether the
user reached checkout
but treatment AFFECTS reaching checkout, and reaching checkout
affects conversion
→ conditioning on checkout-reached manufactures a relationship
between treatment and conversion among a selected subgroup
Never control for anything measured after the treatment. Post-treatment variables are either mediators (removing the effect you want) or colliders (inventing one).
3. Overfitting. With enough predictors and not enough rows, the model fits noise.
40 predictors, 200 observations
in-sample r² 0.78 ← looks excellent
out-of-sample r² 0.11 ← describes nothing
Rule of thumb: at least 10–20 observations per predictor, and validate on data the model didn’t see. With analytics-scale rows this is rarely the binding constraint, but with weekly aggregates — as in Marketing Mix Modelling — it’s the central problem.
Choosing what goes in
From a causal model you can draw, not from what’s in the table.
✓ include: causes of the outcome that also relate to the predictor
(genuine confounders)
pre-treatment characteristics
known drivers, for precision
✗ exclude: anything measured after the treatment
mediators, if you want the total effect
variables highly correlated with another included one
everything else "because we have it"
Stepwise selection — letting an algorithm add and drop predictors by significance — is specifically bad here. It’s a forking-paths machine: the resulting p-values and intervals are invalid because the model was chosen using the same data it’s being tested on.
Reading someone’s model
- Ask what’s not in it. The omitted confounder is where the argument lives
- Check the sample size against the predictor count
- Watch for coefficients that flip sign between model versions — a multicollinearity signature
- Distrust
r²as a quality measure. It rises whenever a predictor is added, whether or not the predictor is meaningful. Adjustedr²penalises this and is the one to quote - Ask whether any predictor is post-treatment
- Remember the coefficient is an association, and no number of controls converts it into a causal effect
Where it interacts
- Linear Regression — the single-predictor case, and where the fitting mechanics are worked
- Confounding Variables — what this is attempting to address and can only partly address
- Propensity Score Matching — a different approach to the same problem, with the same fundamental limitation
- Randomised Controlled Trials — the method that makes all of this unnecessary, where it’s affordable