Testing Strategy
Date: 2026-08-17
Deciding what to test rather than how much. The pyramid is a useful default and a poor rule — the better question is which failures would actually hurt, and what’s the cheapest reliable way to catch each one.
A testing strategy determines what gets tested, at what level, and what deliberately doesn’t.
The pyramid
╱ E2E ╲ few, slow, realistic
╱─────────╲
╱ INTEGRATION ╲ some
╱─────────────────╲
╱ UNIT ╲ many, fast, isolated
The reasoning: unit tests are fast and cheap, so have lots; end-to-end tests are slow and brittle, so have few.
Where it holds: the cost and speed relationship is real, and a suite that’s all end-to-end is slow and flaky.
Where it doesn’t: it says nothing about what to test, and it encourages counting tests rather than considering risk. A codebase can hit any ratio you like while testing none of the things that break.
The criticisms worth knowing
The testing trophy argues integration tests deserve the largest share, because most bugs live at the seams between units rather than inside them — and heavily-mocked unit tests can pass while the assembled system is broken.
The honeycomb makes a similar argument for service-oriented systems.
Both are reacting to the same failure: a codebase with excellent unit coverage where nothing verifies the pieces work together.
UNIT TEST PASSES the function returns
the right shape
INTEGRATION FAILS ...but the API it
calls changed
The pyramid isn’t wrong so much as under-specified. For a typical web application, weighting towards integration is usually the better default.
The question that beats any shape
What failures would actually hurt, and what’s the cheapest way to catch each?
CHECKOUT BREAKS catastrophic
→ E2E, every deploy
VAT CALCULATED WRONG expensive, silent
→ unit tests, many cases
A BUTTON IS THE WRONG minor
COLOUR → visual regression,
or nothing
AN ADMIN PAGE HAS A minor, internal
TYPO → don't test it
Test in proportion to consequence, not to code volume. A payment path deserves more attention than the rest of the application combined; an internal tool used by three people usually deserves very little.
What each level is good at
| Level | Catches | Misses |
|---|---|---|
| Unit | Logic errors, edge cases | Anything about integration |
| Integration | Wiring, contracts, real queries | Full-journey and browser issues |
| E2E | Genuine user journeys | Everything not on the happy path |
| Type checking | Whole classes, at zero runtime cost | Anything about behaviour |
| Manual | Judgement, feel, the unexpected | Regression |
Type checking is the most underrated line in that table. It eliminates a large class of test without anyone writing one — Type Checking in CI.
Coverage, and why it’s a bad target
90% COVERAGE MEANS
90% of lines were EXECUTED during
the test run
IT DOES NOT MEAN
the assertions were meaningful
the edge cases were tested
the code is correct
A test executing code with no assertion counts towards coverage. Coverage measures reach, not verification.
Use it as a discovery tool: low coverage on a critical module is a genuine finding. As a target it produces tests written to hit lines, which are the least useful tests that exist — Goodhart’s law, on schedule.
A workable default for a web application
TYPES strict, everywhere
→ the cheapest tier
UNIT pure logic — pricing,
validation, formatting
→ where the rules are
INTEGRATION the bulk
→ components with real
rendering, API routes
with a real database
E2E a handful of critical
journeys
→ browse → basket →
checkout → confirm
→ run on every deploy
VISUAL the design system
— Visual Regression Testing
MANUAL exploratory, on anything
genuinely new
See: Visual Regression Testing
The critical-journey E2E suite is the one to have first if you have nothing. Half a dozen tests covering the money path catch the failures that matter most, and they justify the rest of the pipeline — End-to-End Testing.
What not to test
Stated explicitly, because it’s rarely written down:
- The framework. It has its own tests
- Third-party libraries
- Trivial getters and setters
- Implementation details — testing internals makes refactoring break tests, which teaches people to avoid refactoring — Unit Testing
- Anything you’d delete rather than fix