Tags: statistics concept

Logistic Regression

Date: 2026-08-17


Regression for yes/no outcomes — converted or not, churned or not, returned or not. It models the log odds rather than the probability, which keeps predictions between 0 and 1 and makes every coefficient an odds ratio: correct, and consistently misreported as a relative change in probability.


Logistic regression predicts the probability of a yes/no outcome from one or more predictors, by fitting a straight line to the log odds of that outcome.

Why not ordinary regression

fitting conversion (0 or 1) with a straight line

ŷ = 0.08 − 0.012 × (form fields)

at 4 fields   ŷ = 0.032   →  3.2%   fine
at 8 fields   ŷ = −0.016  →  −1.6%  ← a negative probability

Linear regression on a binary outcome predicts impossible values, and its errors can’t be constant by construction. The fix is to model something unbounded instead.

Odds and log odds

odds  =  p / (1 − p)

p = 3%    →  odds = 0.03 / 0.97  =  0.0309   ("about 1 in 32")
p = 50%   →  odds = 0.50 / 0.50  =  1.0
p = 80%   →  odds = 0.80 / 0.20  =  4.0

log odds  =  ln(odds)      range: −∞ to +∞    ← now a line can fit it

The model:

ln( p / (1 − p) )  =  a  +  b₁x₁  +  b₂x₂  + …

To get back to a probability:

p  =  1 / (1 + e^−(a + b₁x₁ + …))

Worked

model:  ln(odds of conversion)  =  −3.48  +  0.40 × (returning, 0/1)

NEW CUSTOMER  (returning = 0)
  log odds  =  −3.48
  odds      =  e^−3.48        =  0.0308
  p         =  0.0308 / 1.0308  =  0.0299   →  2.99%

RETURNING     (returning = 1)
  log odds  =  −3.48 + 0.40   =  −3.08
  odds      =  e^−3.08        =  0.0460
  p         =  0.0460 / 1.0460  =  0.0440   →  4.40%

The coefficient as an odds ratio:

odds ratio  =  e^0.40  =  1.492

"returning customers have 1.49× the ODDS of converting"

The misreport that matters

An odds ratio of 1.49 does not mean 49% more likely to convert.

from the worked example
  odds ratio      1.492
  probabilities   2.99%  →  4.40%
  relative risk   4.40 / 2.99  =  1.47      ← close to 1.49, here

At a 3% base rate they’re nearly identical, which is why the error usually goes unnoticed in conversion work. They diverge sharply as the base rate rises:

base rate   odds    × 1.492    new p      relative risk
   3%      0.0309    0.0461     4.41%        1.47   ← OR ≈ RR
  10%      0.1111    0.1658    14.22%        1.42
  30%      0.4286    0.6395    39.01%        1.30
  50%      1.0000    1.4920    59.87%        1.20   ← OR ≫ RR
  80%      4.0000    5.9680    85.65%        1.07

In plain terms: odds ratios approximate relative changes in probability only when the outcome is rare. For conversion rates around 3% you can be loose about it; for email open rates around 40% or checkout completion around 70%, reporting an odds ratio as a percentage lift overstates the effect substantially.

Report predicted probabilities, not coefficients, when talking to anyone who isn’t going to convert log odds in their head. “2.99% versus 4.40%” is unambiguous; “an odds ratio of 1.49” will be repeated as “49% better”.

What it’s used for here

Not usually for analysing an A/B test. A randomised test needs a proportion test, which is simpler, assumption-light and directly interpretable. Reaching for logistic regression on a randomised comparison adds machinery without adding validity — the exception being covariate adjustment for precision, which is a deliberate choice made in advance.

Things that go wrong

  • Separation. If a predictor perfectly splits the outcome, coefficients run off to infinity. Usually means a variable that encodes the outcome — a “reached confirmation page” flag predicting purchase
  • Rare outcomes. With very few positive cases, estimates are biased. A working guide is at least ~10 events per predictor
  • Non-independence. Multiple sessions per user break the assumption and produce standard errors that are too small — same trap as everywhere else — Random Variables
  • Post-treatment predictors. Including something the treatment caused destroys the interpretation — Multiple Regression
  • r² doesn’t apply. Pseudo-r² measures exist and aren’t comparable to linear r²; judge fit by predictive accuracy on held-out data instead

Where it interacts