Tags: experimentation statistics concept
Non-Inferiority Tests
Date: 2026-08-17
A design that asks “is this at least not meaningfully worse”, rather than “is this better”. It’s the right shape for replatforms, cost reductions and simplifications — and it requires committing in advance to how much loss you’d accept, which is the part organisations find uncomfortable and the part that makes it work.
A non-inferiority test tests whether a variant is worse than control by no more than a pre-set margin, rather than whether it’s better.
Why a normal test can’t answer this
The standard test’s null hypothesis is “no difference”. Failing to reject it means you didn’t find a difference — not that there isn’t one.
new checkout, cheaper to maintain. is it as good?
standard A/B test result:
effect −1.2%, 95% CI −4.8% to +2.4%, p = 0.51
read as "no significant difference — ship it"
but the interval includes −4.8%. a 4.8% loss on checkout
is entirely consistent with this data, and would be a disaster
Absence of evidence is not evidence of absence. A wide interval around zero means you learned nothing, and an underpowered test will always deliver one — which is why “we tested it and there was no difference” is so often a statement about sample size — Inconclusive Results.
The design
Flip what has to be proven. You declare a non-inferiority margin (δ) — the largest loss you’re willing to accept — and the test must demonstrate the true effect is better than −δ.
δ = 1% (we'll accept up to a 1% relative loss for the cost saving)
−δ 0
│ │
① ├────────────────┼───────────┤ CI: −0.6% to +1.4%
│ │ entirely above −δ
│ │ → NON-INFERIOR ✓
│ │
② ├──────┼─────────┼───────────┤ CI: −2.1% to +1.8%
│ │ crosses −δ
│ │ → INCONCLUSIVE
│ │
③ ├──┼──┤ │ │ CI: −3.9% to −2.2%
│ │ entirely below −δ
│ │ → INFERIOR ✗
│ │
④ ├────┼───┼───────┤ CI: +0.2% to +2.6%
│ │ above zero
│ │ → SUPERIOR (a bonus)
The decision rule is the position of the interval’s lower bound relative to −δ, and nothing else. Note that case ② — the wide interval — is now correctly labelled inconclusive rather than being mistaken for success, which is the whole gain over the standard test.
Choosing δ
The hard part, and it’s a business decision rather than a statistical one.
Frame it as: what is the change worth, and what loss would that value justify?
new checkout saves £180,000/yr in maintenance and licence costs
checkout revenue is £24,000,000/yr
break-even loss = 180,000 ÷ 24,000,000 = 0.75% of revenue
so a 0.75% conversion loss makes the change value-neutral
→ set δ well inside that. δ = 0.4% leaves real margin
Rules for setting it:
- Set it before launch, in writing, agreed by whoever owns the number. A δ chosen after seeing results is not a margin, it’s a rationalisation — Pre-Registration
- Smaller δ costs more traffic, steeply. Halving δ roughly quadruples the sample needed
- Never set δ larger than the effect a superiority test would have been powered to find. If you’d accept a 3% loss but your normal tests target 3% gains, the margin is wider than your ability to see, and the test is decorative
- It’s a loss you’d accept, not a loss you expect. These are different numbers and people conflate them
Sample size
The arithmetic is the standard one with the margin in place of the effect size — so the sample is driven by δ.
3% baseline, 80% power, 95% confidence (one-sided)
δ = 2.0% relative → ~130,000 per arm
δ = 1.0% relative → ~520,000 per arm
δ = 0.5% relative → ~2,080,000 per arm
In plain terms: proving something isn’t worse by more than a small amount is expensive — often more expensive than proving something is better, because the margins people find acceptable are smaller than the improvements they hope for. This is the practical reason non-inferiority tests get abandoned halfway through.
One-sided is correct here and it’s not a trick: the hypothesis is genuinely directional — you only care about the loss side. Conventionally the one-sided test is run at 2.5% so it corresponds to the lower bound of a 95% two-sided interval, which keeps it comparable with everything else you report.
Where it fits
- Replatforming. New stack, same experience — the question is whether anything was lost — Replatforming, The Strangler Pattern
- Removing something. A feature, a trust badge, a form field. “Does removing this cost us anything” is exactly a non-inferiority question
- Cost reduction. A cheaper image CDN, a lighter recommendation engine, a smaller JavaScript bundle that drops a capability
- Simplification. Fewer steps, fewer options — where you expect a small loss and want to bound it
- Vendor changes. Swapping a payment provider or a search engine
Failure modes
- Running a standard test and reading a null as non-inferiority. The most common error, and the reason this design exists
- δ chosen to make the result work. Check whether δ was recorded before launch; if not, the test is a superiority test with a story
- Ignoring a superiority finding. If the interval sits entirely above zero, you can claim superiority too — that’s legitimate and doesn’t need a separate test
- Guardrails forgotten. Non-inferiority on the primary metric says nothing about the others; the standing guardrail set still applies — Guardrail Metrics
- Treating “inconclusive” as “non-inferior”. Case ② above is a failure to answer the question, and shipping on it is the original mistake wearing better clothes
Where it interacts
- Confidence Intervals — the entire method is reading one interval against one line, which is why interval thinking beats p-value thinking here
- Practical vs Statistical Significance — δ is a practical-significance threshold made explicit and binding
- Statistical Power — the constraint that decides whether the margin you want is affordable
- Post-Test Validation — non-inferiority findings deserve a follow-up check after rollout, since the accepted loss is real and cumulative