Tags: experimentation statistics concept

Non-Inferiority Tests

Date: 2026-08-17


A design that asks “is this at least not meaningfully worse”, rather than “is this better”. It’s the right shape for replatforms, cost reductions and simplifications — and it requires committing in advance to how much loss you’d accept, which is the part organisations find uncomfortable and the part that makes it work.


A non-inferiority test tests whether a variant is worse than control by no more than a pre-set margin, rather than whether it’s better.

Why a normal test can’t answer this

The standard test’s null hypothesis is “no difference”. Failing to reject it means you didn’t find a difference — not that there isn’t one.

new checkout, cheaper to maintain. is it as good?

standard A/B test result:
  effect −1.2%,  95% CI  −4.8% to +2.4%,  p = 0.51

read as "no significant difference — ship it"

but the interval includes −4.8%. a 4.8% loss on checkout
is entirely consistent with this data, and would be a disaster

Absence of evidence is not evidence of absence. A wide interval around zero means you learned nothing, and an underpowered test will always deliver one — which is why “we tested it and there was no difference” is so often a statement about sample size — Inconclusive Results.

The design

Flip what has to be proven. You declare a non-inferiority margin (δ) — the largest loss you’re willing to accept — and the test must demonstrate the true effect is better than −δ.

δ = 1%   (we'll accept up to a 1% relative loss for the cost saving)

                     −δ           0
                      │           │
  ①  ├────────────────┼───────────┤              CI: −0.6% to +1.4%
                      │           │              entirely above −δ
                      │           │              → NON-INFERIOR ✓
                      │           │
  ②  ├──────┼─────────┼───────────┤              CI: −2.1% to +1.8%
                      │           │              crosses −δ
                      │           │              → INCONCLUSIVE
                      │           │
  ③  ├──┼──┤          │           │              CI: −3.9% to −2.2%
                      │           │              entirely below −δ
                      │           │              → INFERIOR ✗
                      │           │
  ④              ├────┼───┼───────┤              CI: +0.2% to +2.6%
                      │           │              above zero
                      │           │              → SUPERIOR (a bonus)

The decision rule is the position of the interval’s lower bound relative to −δ, and nothing else. Note that case ② — the wide interval — is now correctly labelled inconclusive rather than being mistaken for success, which is the whole gain over the standard test.

Choosing δ

The hard part, and it’s a business decision rather than a statistical one.

Frame it as: what is the change worth, and what loss would that value justify?

new checkout saves £180,000/yr in maintenance and licence costs
checkout revenue is £24,000,000/yr

break-even loss  =  180,000 ÷ 24,000,000  =  0.75% of revenue

so a 0.75% conversion loss makes the change value-neutral
→ set δ well inside that. δ = 0.4% leaves real margin

Rules for setting it:

  • Set it before launch, in writing, agreed by whoever owns the number. A δ chosen after seeing results is not a margin, it’s a rationalisation — Pre-Registration
  • Smaller δ costs more traffic, steeply. Halving δ roughly quadruples the sample needed
  • Never set δ larger than the effect a superiority test would have been powered to find. If you’d accept a 3% loss but your normal tests target 3% gains, the margin is wider than your ability to see, and the test is decorative
  • It’s a loss you’d accept, not a loss you expect. These are different numbers and people conflate them

Sample size

The arithmetic is the standard one with the margin in place of the effect size — so the sample is driven by δ.

3% baseline, 80% power, 95% confidence (one-sided)

δ = 2.0%  relative    →   ~130,000 per arm
δ = 1.0%  relative    →   ~520,000 per arm
δ = 0.5%  relative    →  ~2,080,000 per arm

In plain terms: proving something isn’t worse by more than a small amount is expensive — often more expensive than proving something is better, because the margins people find acceptable are smaller than the improvements they hope for. This is the practical reason non-inferiority tests get abandoned halfway through.

One-sided is correct here and it’s not a trick: the hypothesis is genuinely directional — you only care about the loss side. Conventionally the one-sided test is run at 2.5% so it corresponds to the lower bound of a 95% two-sided interval, which keeps it comparable with everything else you report.

Where it fits

  • Replatforming. New stack, same experience — the question is whether anything was lost — Replatforming, The Strangler Pattern
  • Removing something. A feature, a trust badge, a form field. “Does removing this cost us anything” is exactly a non-inferiority question
  • Cost reduction. A cheaper image CDN, a lighter recommendation engine, a smaller JavaScript bundle that drops a capability
  • Simplification. Fewer steps, fewer options — where you expect a small loss and want to bound it
  • Vendor changes. Swapping a payment provider or a search engine

Failure modes

  • Running a standard test and reading a null as non-inferiority. The most common error, and the reason this design exists
  • δ chosen to make the result work. Check whether δ was recorded before launch; if not, the test is a superiority test with a story
  • Ignoring a superiority finding. If the interval sits entirely above zero, you can claim superiority too — that’s legitimate and doesn’t need a separate test
  • Guardrails forgotten. Non-inferiority on the primary metric says nothing about the others; the standing guardrail set still applies — Guardrail Metrics
  • Treating “inconclusive” as “non-inferior”. Case ② above is a failure to answer the question, and shipping on it is the original mistake wearing better clothes

Where it interacts

  • Confidence Intervals — the entire method is reading one interval against one line, which is why interval thinking beats p-value thinking here
  • Practical vs Statistical Significance — δ is a practical-significance threshold made explicit and binding
  • Statistical Power — the constraint that decides whether the margin you want is affordable
  • Post-Test Validation — non-inferiority findings deserve a follow-up check after rollout, since the accepted loss is real and cumulative