Tags: statistics concept

Confounding Variables

Date: 2026-08-17


A confounder is a third variable that causes both the thing you changed and the thing you measured, manufacturing an association between them. It’s the default explanation for any observational finding — and the one class of problem that randomisation solves completely, including for confounders nobody has thought of.


The structure

        C  (confounder)
       ╱ ╲
      ↓   ↓
      A    B

you observe A and B moving together.
neither causes the other. C causes both.

For something to be a confounder it must satisfy three conditions, and checking them is how you argue about a specific case:

1  it causes the OUTCOME (B)
2  it is associated with the EXPOSURE (A)
3  it is NOT on the causal path from A to B

Condition 3 is the one that gets missed. A variable on the path is a mediator, not a confounder — and controlling for a mediator removes the very effect you’re trying to measure.

CONFOUNDER                           MEDIATOR

intent                               ad → click → purchase
 ╱   ╲                                       ↑
↓     ↓                              controlling for "click"
search  purchase                     removes the ad's entire effect
                                     — you've conditioned away the
control for intent → correct         mechanism

Worked

An email campaign appears to work spectacularly:

                 sent email    no email
customers            40,000     160,000
purchased             6,400       9,600
rate                  16.0%        6.0%

apparent effect: +10pp, or 167% relative

Now split by whether they’d purchased in the previous 90 days — the confounder, because recent buyers are both more likely to be on the send list and more likely to buy again:

RECENT BUYERS                  sent      not sent
                             30,000        30,000
purchased                     5,700         5,400
rate                          19.0%         18.0%      +1.0pp

LAPSED                       sent      not sent
                            10,000       130,000
purchased                      700          4,200
rate                          7.0%           3.2%      +3.8pp

within each group: a modest real effect
pooled:            a 167% effect

the campaign was sent disproportionately to people
who were going to buy anyway

In plain terms: 75% of the email list was recent buyers, against 19% of the non-list. The pooled comparison is mostly measuring who was on the list, not what the email did. This is Simpson’s Paradox arising from confounding — and note the effect is real, just five times smaller than claimed.

Why randomisation is different in kind

Adjusting for confounders requires measuring them, knowing about them, and modelling their effect correctly. Randomisation requires none of those.

CONTROLLING FOR CONFOUNDERS          RANDOMISING

works for: the ones you measured     works for: ALL of them —
           and specified correctly              measured, unmeasured,
                                                and unimagined
fails for: everything else
                                     because assignment is independent
"we controlled for device, channel   of every pre-existing characteristic
 and new/returning" — what about     by construction
 intent, mood, urgency, income,
 whether they'd already decided?

This is the single strongest argument for experimentation, and it’s not a matter of degree — it’s the difference between adjusting for a list you wrote and balancing everything simultaneously — Why Randomisation Works, Randomised Controlled Trials.

The two ways adjustment backfires

Controlling for more variables is not safer, and this is counterintuitive enough to be worth stating carefully.

Colliders. A variable caused by both the exposure and the outcome. Controlling for it creates a spurious association where none existed.

        A         B
         ╲       ╱
          ↓     ↓
             C          ← a collider

conditioning on C makes A and B correlated
even when they're causally unrelated

The commerce version: analysing only customers who completed checkout, when both your treatment and the outcome affect completion. Filtering your data is conditioning, so “we only looked at purchasers” can manufacture a relationship — Selection Bias.

Over-adjustment. Controlling for a mediator, as above, removes the real effect and reports nothing.

The rule: choose controls from a causal model you can draw, not from what’s available in the table. Throwing every column into a regression is how colliders and mediators get included — Multiple Regression.

In practice

  • Assume confounding first for any observational comparison. It’s the most likely explanation and the one to argue against explicitly
  • Name the specific confounder rather than objecting in general. “Purchase intent” is a claim someone can test; “correlation isn’t causation” isn’t
  • Check pre-period differences. If two groups differed before the intervention, they differ for a reason — Difference-in-Differences
  • Randomise wherever you can afford to. Everything else on this page is a workaround for not having done it
  • Where you can’t, be explicit about what remains unadjusted. “Controlled for device and channel; not for intent, which we can’t observe” is an honest statement of what the number is worth — Propensity Score Matching

Where it interacts