End-to-End Testing
Date: 2026-08-17
Driving a real browser through a real user journey against a running system. It’s the only tier that proves the thing actually works — and it’s slow and flaky enough that the discipline is having very few, covering only what would be catastrophic.
An end-to-end (E2E) test automates a browser through a complete user journey against a deployed or locally-running application, with real network, real rendering and usually a real database.
What only this tier catches
the CSS that hides the button on mobile
the third-party script that blocks
the page — Third-Party Scripts
the redirect that loses the basket
the cookie banner covering "Continue"
the payment iframe that won't accept
focus — Focus Management
the CDN serving a stale bundle
the build that succeeded and shipped
nothing
See: Third-Party Scripts · Focus Management
Every one of those passes every unit and integration test. They’re integration failures between systems nobody owns together — which is precisely the gap this tier exists for.
What to test
Very few journeys, chosen by consequence:
ESSENTIAL
browse → product → basket →
checkout → payment → confirmation
login → account → reorder
search → filter → product
THAT'S ROUGHLY IT
If the money path works, the site works enough to trade. Everything else belongs at a cheaper tier — Testing Strategy.
The failure mode is scope creep. An E2E suite that grows to 300 tests becomes slow, flaky and eventually ignored, and its failures stop being investigated — which is worse than not having it.
Playwright, in practice
test('completes a purchase', async ({ page }) => {
await page.goto('/products/serum')
await page.getByRole('button',
{ name: 'Add to basket' }).click()
await page.getByRole('link',
{ name: 'Checkout' }).click()
await page.getByLabel('Email')
.fill('test@example.com')
await page.getByLabel('Postcode')
.fill('SW1A 1AA')
await page.getByRole('button',
{ name: 'Place order' }).click()
await expect(page.getByRole('heading',
{ name: /order confirmed/i })).toBeVisible()
})Two things carry the reliability here:
- Role and label selectors, not CSS classes — resilient to styling changes, and they verify accessibility incidentally
- Web-first assertions.
expect(...).toBeVisible()retries until a timeout rather than checking once, which removes most manual waiting
Never use fixed waits:
await page.waitForTimeout(2000) // never
await expect(locator).toBeVisible() // yeswaitForTimeout is the single largest cause of flaky E2E suites — too short and it fails intermittently, too long and the suite crawls — Flaky Tests.
Where to run them
LOCAL on demand, while developing
PULL REQUEST the critical journeys only
→ against a preview
environment
— Preview Environments
POST-DEPLOY the same suite against
production
→ a smoke test that the
deploy actually works
→ the highest-value run
See: Preview Environments
Running the critical journey against production after every deploy is the most valuable E2E run there is. It catches the environment-specific failures that no pre-production test can — and it’s the difference between finding out in ninety seconds and finding out from a customer.
Test data and state
The hardest part of this tier, and the usual reason suites become unreliable:
- Seed deterministically. A test depending on whatever data happens to exist is a test that fails on a Tuesday
- Create what you need per test, via an API rather than through the UI — driving the browser through setup is slow and couples the test to unrelated screens
- Clean up, or use disposable data. Namespaced accounts, or a reset between runs
- Never test against shared mutable state. Two runs in parallel modifying the same order is a race — Race Conditions
Payments
SANDBOX MODE the provider's test
environment and test cards
← use this
MOCK THE GATEWAY faster, less realistic
← reasonable for PR runs
REAL PAYMENTS never
Sandbox mode is what it’s for, and every provider has one. Test declines and timeouts as well as successes — the decline path is where the interesting failures are and it’s rarely covered.
The honest limitations
- Slow. Seconds per test, minutes per suite
- Flaky by nature. Real network, real timing, real browser
- Failures don’t localise. “Checkout failed” could be twenty things
- Expensive to maintain. UI changes break them, legitimately
Which is the argument for having few. Six reliable tests that always run and are always investigated beat two hundred that people have learned to re-run until green.