Flaky Tests
Date: 2026-08-17
A test that passes and fails on the same code. One is an annoyance; a handful destroys the suite’s value entirely, because people start re-running failures instead of reading them — and then a real failure gets re-run too.
A flaky test produces different results on identical code. It’s non-deterministic, and the non-determinism is nearly always in the test rather than in the application.
Why one is not a small problem
1 flaky test, 99% reliable
→ a 200-test suite fails ~1 run in 2
CONSEQUENCE
"just re-run it" becomes the habit
↓
a REAL failure gets re-run too
↓
the suite no longer gates anything
The cost isn’t the flaky test — it’s the behaviour it teaches. Once re-running is normal, the suite has stopped being a signal and become a toll.
The causes
Timing and race conditions — the largest category by a distance:
// flaky — is it there yet?
await page.click('#submit')
expect(page.locator('.success')).toBeVisible()
// stable — retries until it is
await page.getByRole('button',
{ name: 'Submit' }).click()
await expect(
page.getByText('Order confirmed')).toBeVisible()Fixed waits are the specific culprit:
await page.waitForTimeout(2000) // neverToo short and it fails on a slow day; too long and the suite crawls. Wait for a condition, never for a duration — End-to-End Testing.
Shared state — tests interfering through the database, a module-level variable, a cache, or the filesystem. Presents as “passes alone, fails in the suite” — Test Data.
Order dependence — test B relies on something test A created. Invisible until the runner parallelises or shuffles.
Time — a test asserting on “today” that breaks at midnight, at month end, or when CI runs in UTC and you don’t.
Randomness — unseeded random data occasionally generating a value that breaks an assumption. This one is a gift: it’s usually found a real bug.
External dependencies — a third-party API being slow or rate-limiting. Fix by not calling it — Test Doubles.
Resource contention — parallel workers on a loaded CI machine, where everything is slower and timing assumptions fail.
Finding them
# does it survive repetition?
vitest --repeat 50 path/to/test
# does it survive a different order?
vitest --sequence.shuffleTrack flakiness in CI rather than relying on memory. Most CI platforms can flag tests that failed and then passed on retry — that list is the backlog, and without it the same three tests annoy everyone forever without anyone owning them.
The policy that works
1 DETECT — track pass-after-retry
2 QUARANTINE — move it out of the
blocking suite immediately
3 TICKET IT — with an owner and a date
4 FIX or DELETE within the window
Quarantine, not tolerate. Leaving a known-flaky test in the blocking suite is what trains the re-run habit. Removing it from the gate keeps the signal clean while it’s fixed.
And enforce the deadline. A quarantine with no expiry is deletion with extra steps — which is fine, as long as it’s an explicit decision.
Retries
retries: 2Retries hide flakiness rather than fixing it, and used unconditionally they let the suite rot invisibly.
Where they’re defensible: end-to-end tests against real infrastructure, where some non-determinism is genuinely environmental. Even then, log every retry — a rising retry count is the metric that says the suite is degrading.
Never retry unit or integration tests. Those should be deterministic, and a flaky one is a bug.
Sometimes it’s the application
The case worth checking before blaming the test:
a flaky test can be finding a REAL
race condition
→ two requests, order-dependent
→ an unawaited promise
→ a genuine timing bug users hit
occasionally
Before quarantining, ask whether the non-determinism is in the code under test. An intermittent failure that only happens under load is exactly what a production race looks like — Race Conditions.