Tags: experimentation concept
A-A Tests
Date: 2026-08-16
Run the identical experience against itself and check the platform reports nothing. It’s the only test whose correct answer you already know, which makes it the only one that can validate the machinery rather than the change.
What it is
An A/A test is an experiment where both arms receive the same experience. Assignment, tracking, and analysis run exactly as they would for a real test, and the expected result is no significant difference.
Any significance it produces beyond the expected rate is a fault in the system, not a finding about users.
What it catches
The failures that would otherwise silently corrupt every subsequent test:
- Sample Ratio Mismatch — the split doesn’t land where it should, so assignment or delivery is broken
- Asymmetric tracking — one arm’s events fire more reliably than the other’s, usually because the variant path loads extra code
- Non-uniform bucketing — the hash correlates with something about the identifier, so the arms differ before anything is changed. See Assignment and Bucketing
- Assignment leakage — users appearing in both arms, or reassigned on login
- A miscalibrated significance calculation — the tool’s maths is wrong, or it’s treating sessions as independent when they aren’t
- Contamination — bots or internal traffic hitting one arm preferentially, see Bot and Internal Traffic
In plain terms: you cannot tell whether a testing tool is lying to you by running tests where you don’t know the answer. An A/A test is the one case where you do.
What “expected rate” means
The unintuitive part. At 95% significance, an A/A test should come back significant 5% of the time. That’s the definition of the threshold, not a fault.
20 A/A tests at α = 0.05
expected false positives = 20 × 0.05 = 1
So one significant A/A result is not evidence of a broken platform. Four out of twenty is. This is the same arithmetic as The Multiple Comparisons Problem, applied to your own quality assurance.
The practical implication: a single A/A test is weak evidence. It catches gross failures — SRM, obvious tracking asymmetry — and cannot detect subtler calibration problems. For those you need either many A/A tests or a simulation over historical data.
The stronger version
Rather than running A/A tests in production and waiting weeks, replay historical data through the analysis pipeline with randomly assigned fake arms, hundreds of times. Then check that the distribution of p-values is roughly uniform and that significance appears at about 5%.
This catches calibration errors a handful of live A/A tests never would, costs no traffic, and takes minutes. It’s the check worth automating.
When to run one
- Before trusting a new platform. Non-negotiable, and the results are worth keeping
- After changing the assignment mechanism, the tracking implementation, or the analysis code
- After a replatform — new front end, new CDN, new tag setup
- Alongside a real test, as a third arm. Cheap insurance, at the cost of some traffic
- When results look too good. A run of large wins is more often a broken platform than a brilliant quarter
Cost and limits
An A/A test consumes exactly as much traffic as a real one and produces no business value, which is why they’re skipped. On a nine-week testing cadence that’s a real sacrifice, and the honest mitigation is the simulation approach above plus continuous SRM monitoring rather than periodic full A/A runs.
Limits worth knowing:
- It validates the pipeline, not the design. A perfectly calibrated platform will still mislead you if the Randomisation Unit is wrong or the metric is badly defined
- It can’t detect problems that only appear when the variant differs — Flicker and Flash of Original Content by definition doesn’t occur when both arms are identical, which is precisely why flicker is invisible to A/A validation
- It says nothing about power. A clean A/A test on an underpowered design still gives you an underpowered design