Tags: web-dev concept

Flaky Tests

Date: 2026-08-17


A test that passes and fails on the same code. One is an annoyance; a handful destroys the suite’s value entirely, because people start re-running failures instead of reading them — and then a real failure gets re-run too.


A flaky test produces different results on identical code. It’s non-deterministic, and the non-determinism is nearly always in the test rather than in the application.

Why one is not a small problem

1 flaky test, 99% reliable
  → a 200-test suite fails ~1 run in 2

CONSEQUENCE
  "just re-run it" becomes the habit
  ↓
  a REAL failure gets re-run too
  ↓
  the suite no longer gates anything

The cost isn’t the flaky test — it’s the behaviour it teaches. Once re-running is normal, the suite has stopped being a signal and become a toll.

The causes

Timing and race conditions — the largest category by a distance:

// flaky — is it there yet?
await page.click('#submit')
expect(page.locator('.success')).toBeVisible()
 
// stable — retries until it is
await page.getByRole('button',
  { name: 'Submit' }).click()
await expect(
  page.getByText('Order confirmed')).toBeVisible()

Fixed waits are the specific culprit:

await page.waitForTimeout(2000)   // never

Too short and it fails on a slow day; too long and the suite crawls. Wait for a condition, never for a duration — End-to-End Testing.

Shared state — tests interfering through the database, a module-level variable, a cache, or the filesystem. Presents as “passes alone, fails in the suite” — Test Data.

Order dependence — test B relies on something test A created. Invisible until the runner parallelises or shuffles.

Time — a test asserting on “today” that breaks at midnight, at month end, or when CI runs in UTC and you don’t.

Randomness — unseeded random data occasionally generating a value that breaks an assumption. This one is a gift: it’s usually found a real bug.

External dependencies — a third-party API being slow or rate-limiting. Fix by not calling it — Test Doubles.

Resource contention — parallel workers on a loaded CI machine, where everything is slower and timing assumptions fail.

Finding them

# does it survive repetition?
vitest --repeat 50 path/to/test
 
# does it survive a different order?
vitest --sequence.shuffle

Track flakiness in CI rather than relying on memory. Most CI platforms can flag tests that failed and then passed on retry — that list is the backlog, and without it the same three tests annoy everyone forever without anyone owning them.

The policy that works

1  DETECT — track pass-after-retry
2  QUARANTINE — move it out of the
   blocking suite immediately
3  TICKET IT — with an owner and a date
4  FIX or DELETE within the window

Quarantine, not tolerate. Leaving a known-flaky test in the blocking suite is what trains the re-run habit. Removing it from the gate keeps the signal clean while it’s fixed.

And enforce the deadline. A quarantine with no expiry is deletion with extra steps — which is fine, as long as it’s an explicit decision.

Retries

retries: 2

Retries hide flakiness rather than fixing it, and used unconditionally they let the suite rot invisibly.

Where they’re defensible: end-to-end tests against real infrastructure, where some non-determinism is genuinely environmental. Even then, log every retry — a rising retry count is the metric that says the suite is degrading.

Never retry unit or integration tests. Those should be deterministic, and a flaky one is a bug.

Sometimes it’s the application

The case worth checking before blaming the test:

a flaky test can be finding a REAL
race condition
  → two requests, order-dependent
  → an unawaited promise
  → a genuine timing bug users hit
    occasionally

Before quarantining, ask whether the non-determinism is in the code under test. An intermittent failure that only happens under load is exactly what a production race looks like — Race Conditions.