Tags: statistics concept
One-Tailed vs Two-Tailed Tests
Date: 2026-08-17
Whether you’re testing for a difference in either direction or only in one. A one-tailed test roughly halves your p-value for free, which is exactly why it’s so often reached for — and why choosing it after seeing which way the result went is one of the cleanest forms of P-Hacking available.
A two-tailed test counts a surprising result in either direction as evidence; a one-tailed test counts only one direction, decided before the data is seen.
The mechanics
The test statistic is the same. What changes is which part of the distribution counts as “surprising”.
TWO-TAILED ONE-TAILED (upper)
▁▂▃▅▇█▇▅▃▂▁ ▁▂▃▅▇█▇▅▃▂▁
███ ███ █████
2.5% 2.5% 5%
└─ reject ─┘ └ reject ┘
α split across both ends all of α at one end
critical z = ±1.96 critical z = 1.645
asks: "is B different from A?" asks: "is B better than A?"
a harmful result is
indistinguishable from no effect
Worked, on the same data:
observed z = 1.78
two-tailed p = 2 × P(Z > 1.78) = 2 × 0.0375 = 0.075 → not significant
one-tailed p = P(Z > 1.78) = 0.0375 → "significant"
In plain terms: the identical result flips from “no conclusion” to “ship it” purely because of a choice about which question you claimed to be asking. Nothing about the data changed.
Why it’s almost always wrong here
1. You do care about the other direction. A one-tailed upper test treats a catastrophic result exactly like a null one — both land in “fail to reject”. If checkout conversion drops 4%, a one-tailed test for improvement reports the same “not significant” it would report for no change.
result: variant is 4% WORSE, z = −2.4
two-tailed p = 0.016 → significant. it's harmful. don't ship,
and now you've learned something
one-tailed p = 0.992 → "not significant"
→ reads as "no difference, ship if you like"
That’s not a technicality — it’s the whole reason two-tailed is the default. You are never indifferent to harm in commercial testing.
2. It must be chosen before the data. The direction has to be pre-registered, with a justification that doesn’t depend on the outcome. In practice the sequence is: run the test, see 0.075, remember one-tailed exists, report 0.0375. That’s an increase in the false positive rate from 5% to effectively 10%, and it’s invisible in the write-up — Pre-Registration, The Garden of Forking Paths.
3. The power argument is smaller than it sounds. The genuine benefit is real but modest:
detecting a 5% relative lift, 3% baseline, 80% power, 95% confidence
two-tailed ≈ 210,000 per arm
one-tailed ≈ 165,000 per arm ~21% less traffic
A fifth less traffic, in exchange for being unable to detect harm. For a change that could plausibly hurt — which is most changes — that’s a bad trade.
When it’s legitimate
Narrow, and both cases share a property: the negative direction leads to the same action as the null.
- Non-Inferiority Tests. The whole design is directional — you’re establishing the effect is better than −δ, and “much better” and “slightly better” lead to the same decision. This is the one common legitimate use in commerce, and it’s conventionally run at a one-sided 2.5% so it corresponds to the lower bound of a 95% two-sided interval
- A physically impossible direction. Rare and usually an illusion — “removing a step can’t reduce completion” is exactly the assumption that gets falsified
- Regulated superiority claims, where the framework specifies it
Note the convention that keeps everything comparable: running one-sided at 2.5% rather than 5% gives you the directional question without the free significance boost. If someone proposes one-tailed for the power saving, this is the counter-offer — it removes the incentive and keeps the arithmetic honest.
Reading someone else’s result
The check to apply when handed a p-value:
- Was the tail choice declared before launch? If the analysis plan doesn’t say, assume two-tailed and double the p-value you were given
- Would a negative result have been reported? If a 4% drop would have been written up as “inconclusive”, the test was one-tailed in effect regardless of what was run
- Does the tool default to one-tailed? Some do. Check rather than assume, because it changes every result the tool has ever reported to you
- Prefer the interval. A confidence interval sidesteps the whole argument — it shows the plausible range in both directions, and you can see harm and benefit at once — Confidence Intervals
Where it interacts
- Hypothesis Testing — the frame this is a parameter of, and where the null gets stated
- Null and Alternative Hypotheses — the tail choice is a property of how the alternative is written, decided at design time
- Statistical Power — the 21% saving above, and why it isn’t worth what it costs
- Reading a Test Result — where in the ordered checks this gets verified