Tags: experimentation statistics concept
Inconclusive Results
Date: 2026-08-16
The most common outcome, and not a failure. “Not significant” collapses three genuinely different situations into one word — and the confidence interval is what tells them apart.
What it is
An inconclusive result is one where the data cannot distinguish between the outcomes you care about. It is not a finding of no effect, and it is not a failed test.
The three situations it hides:
CI on relative lift reading
−0.4% to +0.6% ●─● precise null. Ruled out anything meaningful.
A real, useful finding
−3.9% to +9.9% ●─────● uninformative. Could be a solid win or a
solid loss. You learned nothing
+1.8% to +7.4% ●──● probably a win, but the lower bound sits
below the £15k threshold. Underpowered,
not negative
All three report as “not significant”. Only the first is evidence of no effect; the third is closer to a win than a loss.
In plain terms: the p-value tells you whether zero is plausible. The interval tells you what else is plausible, and that’s the part you can act on.
Reading one properly
Three questions, in order:
- What does the interval rule out? If the upper bound is below your commercial threshold, the change genuinely isn’t worth shipping — that’s a decision, not an absence of one
- What was the MDE? “No significant difference” from a test that could only have detected a 15% lift says nothing about a 5% one. Always report the MDE alongside a null result
- Where does the point estimate sit? A +4% point estimate with a wide interval is weak evidence for a win. Weak evidence isn’t no evidence, and for a cheap reversible change it may be enough
Why they dominate
Most tests come back inconclusive, and the reasons are structural rather than a sign of poor ideas:
- Most true effects are small — smaller than the MDE most sites can afford
- Most sites are traffic-constrained. Nine weeks for a 10% MDE at 3% conversion is the norm, not an outlier — Test Duration
- Timid changes have timid effects. The traffic cost of detecting a small change is enormous, because sample scales with the square of the effect
A realistic programme sees roughly one in five to one in eight tests produce a clear, shippable win. A programme reporting a much higher win rate is more likely miscalibrated than brilliant — check for Peeking, multiple metrics, and post-hoc segmentation.
What to do with one
- Record it fully. Hypothesis, mechanism, interval, MDE. An inconclusive result with a mechanism narrows the space of future ideas — Experiment Archive, Institutional Learning
- Decide whether to retest. Worth it if the point estimate is promising and you can materially raise power: bolder change, wider exposure, Variance Reduction, Winsorisation and Capping. Rerunning the identical design gets you the same answer
- Don’t rerun until it wins. Repeating a test until it crosses the line is The Multiple Comparisons Problem spread over months, and it works — which is the problem
- Consider shipping anyway. For a cheap reversible change with a positive point estimate and clean guardrails, shipping on weak evidence is often the right call. Say that’s what you’re doing rather than reporting it as a win
- Consider not building the next one. If the change was expensive and the interval spans zero, the honest answer may be that this class of change can’t be validated at your traffic — Rollouts as Experiments and Holdout Groups are the alternatives
How to report it
Never as “the test failed”. A defensible form:
“No detectable difference. The 95% interval runs from −0.4% to +0.6%, so we can rule out anything above roughly half a percent in either direction — this change does not move conversion. The mechanism we proposed, that delivery cost drives abandonment at this step, is not supported.”
That’s a result. Compare with “inconclusive, we’ll try something else”, which records nothing and leaves the same idea available to be re-proposed next year.
See Communicating Uncertainty for the wider problem of presenting ranges to audiences that want a number.