Tags: statistics concept
Survivorship Bias
Date: 2026-08-16
Measuring only what’s still there to measure. The failures left the dataset, so the analysis describes survivors and gets read as describing everyone — and the missing cases are usually the informative ones.
What it is
Survivorship bias is a form of Selection Bias where the sample contains only units that persisted long enough to be observed, and the ones that didn’t are invisible rather than recorded as absent.
The distinguishing feature: the missing data doesn’t look missing. A dataset of survivors is complete, internally consistent, and describes a population that isn’t the one you meant.
The canonical illustration
Wartime analysts examined returning aircraft to decide where to add armour, and found bullet holes concentrated on the wings and fuselage. The proposal was to reinforce those areas.
The correction: those are the places a plane can be hit and still return. The planes hit in the engines weren’t in the sample, because they didn’t come back. Armour belonged where the returning planes showed no damage.
The structure recurs everywhere: the data shows you the survivors, and the answer is in the absences.
Where it appears in your work
- Retention analysis on identified users. Anyone who never returned far enough to log in is absent. Your retention curve describes people who came back at least once — Retention Curves, Anonymous and Identified Users
- Session replay. You watch sessions that were recorded. Users who left before the recorder loaded, or who blocked it, aren’t there — and they’re disproportionately the ones with a bad experience — Session Replay
- Real user monitoring. Someone who abandoned during a slow load may never send a beacon. Your performance data is systematically optimistic, and you cannot fix it — only know it — Real User Monitoring
- Survey responses. Answered by people still engaged enough to answer
- Customer feedback and reviews. From those who stayed, plus a vocal minority who left angry. The silent majority who drifted away are absent
- Experiment archives that only record winners. A programme that files wins and forgets losses cannot compute its own win rate and will overestimate itself indefinitely — Experiment Archive
- “Successful companies do X” reasoning. The companies that did X and failed aren’t in the book
The test
One question, applied to any dataset:
What would be missing from this if it happened?
If the answer is “the thing I care about”, you have survivorship bias. A retention dataset can’t contain people who never came back. A replay dataset can’t contain sessions that never recorded. A performance dataset can’t contain loads that were abandoned.
What to do
- Find the denominator elsewhere. Reconcile against a system that sees everyone — the order system, server logs, the CRM. The gap between “everyone who arrived” and “everyone in the analytics” is the size of what’s missing — Guide - Auditing a Tracking Plan
- Measure at the earliest possible point. A metric recorded at page start survives more than one recorded at page end
- State the direction of the bias even when you can’t quantify it. “This retention curve excludes users who never authenticated, so it’s optimistic” is a useful sentence
- Record failures deliberately. Losses and inconclusives in the experiment archive, abandoned sessions in the funnel, churned customers in the cohort — the absences have to be constructed, because they won’t arrive on their own
Why it’s hard to notice
The dataset is complete and coherent. Nothing errors, nothing is flagged, and every number is internally consistent. There’s no signal to notice — which is why it has to be found by reasoning about the collection mechanism rather than by inspecting the data.
Related: Regression to the Mean, where selecting on an extreme produces a spurious improvement — the same shape, applied to selecting on a value rather than on survival.