Tags: statistics concept

Multiple Regression

Date: 2026-08-17


Fitting several predictors at once, so each coefficient is the effect of that variable “holding the others constant”. The phrase is what makes it dangerous — it sounds like experimental control and isn’t, because it only holds constant the things you measured and put in the model.


Multiple regression is linear regression with more than one predictor, estimating each predictor’s slope while the others are held fixed in the model.

The model

y  =  a  +  b₁x₁  +  b₂x₂  +  b₃x₃  +  error

Each b is the change in y per unit of that x, with the other predictors held fixed.

conversion rate  =  4.9  −  0.31·(load seconds)
                        +  1.42·(returning customer, 0/1)
                        −  0.08·(form fields)

b₁ = −0.31   each extra second costs 0.31pp, among customers
             of the same returning-status and form length
b₂ = +1.42   returning customers convert 1.42pp higher, at the
             same load time and form length
b₃ = −0.08   each extra form field costs 0.08pp, all else equal

Note that b₁ fell from −0.44 in the simple regression to −0.31 here. The difference is what the other variables were absorbing — part of what looked like a load-time effect was actually returning customers being served faster pages.

”Controlling for” is not control

The central point of the note.

EXPERIMENTAL CONTROL                 STATISTICAL CONTROL

randomise                            add a column to a regression

balances EVERY variable —            adjusts for the variables you
measured, unmeasured, and            measured, specified correctly,
never imagined                       and modelled with the right
                                     functional form

— Why Randomisation Works        everything else remains confounded

Randomisation achieves the left-hand column’s guarantee without needing the list at all — Why Randomisation Works.

In plain terms: a regression controlling for device, channel and returning-status has not controlled for purchase intent, urgency, income, whether the customer had already decided, or anything else you don’t have a column for. Those are usually the variables that matter most, and they’re unmeasured precisely because they’re hard to measure — Confounding Variables.

The language does real damage here. “We controlled for X” is heard as “X has been eliminated as an explanation”, and what it means is “X is one of the several dozen possible explanations we were able to adjust for”.

Three ways adding variables makes it worse

1. Multicollinearity. Predictors that move together can’t be separated.

model includes both "page weight (KB)" and "number of images"
correlation between them: 0.94

fitted coefficients

  page weight    −0.002  (SE 0.004)   ← "not significant"
  images         −0.140  (SE 0.210)   ← "not significant"

drop one, and the other becomes strongly significant.

the model cannot attribute the effect between two variables
that always move together. it isn't confused — the DATA
contains no information to separate them

Symptoms: large standard errors, coefficients that flip sign when a variable is added or removed, a model with high overall r² and no individually significant predictors.

2. Colliders. Adding a variable caused by both the predictor and the outcome creates a spurious association that wasn’t there.

regressing conversion on treatment, "controlling for" whether the
user reached checkout

but treatment AFFECTS reaching checkout, and reaching checkout
affects conversion

→ conditioning on checkout-reached manufactures a relationship
  between treatment and conversion among a selected subgroup

Never control for anything measured after the treatment. Post-treatment variables are either mediators (removing the effect you want) or colliders (inventing one).

3. Overfitting. With enough predictors and not enough rows, the model fits noise.

40 predictors, 200 observations

in-sample r²      0.78     ← looks excellent
out-of-sample r²  0.11     ← describes nothing

Rule of thumb: at least 10–20 observations per predictor, and validate on data the model didn’t see. With analytics-scale rows this is rarely the binding constraint, but with weekly aggregates — as in Marketing Mix Modelling — it’s the central problem.

Choosing what goes in

From a causal model you can draw, not from what’s in the table.

✓  include:  causes of the outcome that also relate to the predictor
             (genuine confounders)
             pre-treatment characteristics
             known drivers, for precision

✗  exclude:  anything measured after the treatment
             mediators, if you want the total effect
             variables highly correlated with another included one
             everything else "because we have it"

Stepwise selection — letting an algorithm add and drop predictors by significance — is specifically bad here. It’s a forking-paths machine: the resulting p-values and intervals are invalid because the model was chosen using the same data it’s being tested on.

Reading someone’s model

  • Ask what’s not in it. The omitted confounder is where the argument lives
  • Check the sample size against the predictor count
  • Watch for coefficients that flip sign between model versions — a multicollinearity signature
  • Distrust r² as a quality measure. It rises whenever a predictor is added, whether or not the predictor is meaningful. Adjusted r² penalises this and is the one to quote
  • Ask whether any predictor is post-treatment
  • Remember the coefficient is an association, and no number of controls converts it into a causal effect

Where it interacts