Tags: statistics experimentation concept

Proportion Tests

Date: 2026-08-16


Comparing two rates. It’s the workhorse — nearly every conversion test is this — and the chi-squared and z-test forms are the same question asked twice, which is why tools disagree on the label and agree on the answer.


What it is

A proportion test compares the rate of a binary outcome between two groups: did the conversion rate differ between control and variant?

The mechanism, in three steps:

1  pool the two arms to estimate the rate under the null
     (the null: the assumption that there's no real difference)
2  compute the standard error of the difference
3  express the observed difference in standard errors

Worked, on 53,000 per arm:

control    1,590 / 53,000  =  3.000%
variant    1,638 / 53,000  =  3.090%
difference                 =  0.090 pp

pooled p̄  = (1,590 + 1,638) / 106,000        = 0.03045

SE        = √[2 × 0.03045 × 0.96955 ÷ 53,000]
          = √0.000001114
          = 0.001056  →  0.1056 pp

z         = 0.090 ÷ 0.1056                    = 0.85
p         = 0.395  (two-tailed)

Pooled p̄ is both arms combined into one rate — the best guess at the true rate if the null is right. Standard error (SE) is how much the difference between two arms would wobble by chance alone at this sample size. z is the observed difference measured in those wobble-units; the p-value is how often chance alone produces a z at least that big.

Not significant. Under the null you’d see a gap this large or larger about 40% of the time.

In plain terms: the variant did a bit better, but two identical pages split this way would show a gap this size two times in five. Nothing here separates the result from luck.

z-test and chi-squared are the same test

Genuinely, not approximately. For a 2×2 comparison:

z    = 0.85
z²   = 0.72  =  χ²

Same p-value, different arithmetic route. So a tool reporting “chi-squared” and one reporting “z-test for proportions” will agree, and the label tells you nothing about the method’s quality.

Chi-squared generalises to more than two groups or more than two outcomes; the z-test gives you a signed direction and a confidence interval directly, which is usually what you want. Prefer the z-test framing for A/B tests for that reason.

The confidence interval

More useful than the p-value, and it comes from the same standard error — though the interval conventionally uses the unpooled standard error rather than the pooled one used for the test:

95% CI on the difference
  = 0.090 ± 1.96 × 0.1056
  = −0.117 pp  to  +0.297 pp

as a relative lift on a 3% base
  = −3.9%  to  +9.9%

The unpooled SE, √[0.03 × 0.97 ÷ 53,000 + 0.0309 × 0.9691 ÷ 53,000], comes out at 0.1055 pp — the same to three figures here, because the two rates are close. They diverge when the arms differ a lot.

In plain terms: the true effect could plausibly be anything from a 4% drop to a 10% gain. That range includes zero, which is the same “can’t tell” verdict as the p-value, but it also shows how big the possible upside and downside are.

Report this, not the p-value. It carries magnitude and precision; the p-value carries neither — Confidence Intervals, Effect Size.

When it’s valid

  • Independent observations. One row per unit, each counted once. Sessions as the unit when users return violates this and overstates significance — Randomisation Unit
  • Enough successes and failures. Roughly 10 expected in each cell of each arm. At very low base rates the normal approximation degrades and an exact method is needed — Binomial and Bernoulli Distributions
  • Fixed sample, evaluated once. Checking repeatedly invalidates it — Peeking
  • Random assignment — otherwise you’re measuring a difference between populations, not an effect

When to reach for something else

SituationInstead
Continuous metric — revenue, order valuet-test, or better Bootstrapping
More than two variantsChi-squared, then pairwise with a correction — The Multiple Comparisons Problem
Very rare outcomeExact test
Ratio metric with a varying denominatorRatio Metrics — the naive standard error is wrong
Continuous monitoring requiredSequential Testing, Always-Valid Inference

Practical

  • Don’t hand-compute it. Every tool does this correctly; the value is knowing what it did
  • Check the inputs before the output. Both arms’ denominators should match the intended split — if they don’t, the test is invalid regardless of the result — Sample Ratio Mismatch
  • Read the interval, convert it to money, then apply your decision rule — Guide - Statistics for CRO
  • One test, one primary metric. Running it on six metrics is six comparisons, whatever the tool calls it