Tags: web-dev concept

Testing Strategy

Date: 2026-08-17


Deciding what to test rather than how much. The pyramid is a useful default and a poor rule — the better question is which failures would actually hurt, and what’s the cheapest reliable way to catch each one.


A testing strategy determines what gets tested, at what level, and what deliberately doesn’t.

The pyramid

        ╱ E2E ╲          few, slow, realistic
      ╱─────────╲
    ╱ INTEGRATION ╲      some
  ╱─────────────────╲
╱       UNIT         ╲   many, fast, isolated

The reasoning: unit tests are fast and cheap, so have lots; end-to-end tests are slow and brittle, so have few.

Where it holds: the cost and speed relationship is real, and a suite that’s all end-to-end is slow and flaky.

Where it doesn’t: it says nothing about what to test, and it encourages counting tests rather than considering risk. A codebase can hit any ratio you like while testing none of the things that break.

The criticisms worth knowing

The testing trophy argues integration tests deserve the largest share, because most bugs live at the seams between units rather than inside them — and heavily-mocked unit tests can pass while the assembled system is broken.

The honeycomb makes a similar argument for service-oriented systems.

Both are reacting to the same failure: a codebase with excellent unit coverage where nothing verifies the pieces work together.

UNIT TEST PASSES     the function returns
                     the right shape

INTEGRATION FAILS    ...but the API it
                     calls changed

The pyramid isn’t wrong so much as under-specified. For a typical web application, weighting towards integration is usually the better default.

The question that beats any shape

What failures would actually hurt, and what’s the cheapest way to catch each?

CHECKOUT BREAKS          catastrophic
                         → E2E, every deploy

VAT CALCULATED WRONG     expensive, silent
                         → unit tests, many cases

A BUTTON IS THE WRONG    minor
COLOUR                   → visual regression,
                           or nothing

AN ADMIN PAGE HAS A      minor, internal
TYPO                     → don't test it

Test in proportion to consequence, not to code volume. A payment path deserves more attention than the rest of the application combined; an internal tool used by three people usually deserves very little.

What each level is good at

LevelCatchesMisses
UnitLogic errors, edge casesAnything about integration
IntegrationWiring, contracts, real queriesFull-journey and browser issues
E2EGenuine user journeysEverything not on the happy path
Type checkingWhole classes, at zero runtime costAnything about behaviour
ManualJudgement, feel, the unexpectedRegression

Type checking is the most underrated line in that table. It eliminates a large class of test without anyone writing one — Type Checking in CI.

Coverage, and why it’s a bad target

90% COVERAGE MEANS
  90% of lines were EXECUTED during
  the test run

IT DOES NOT MEAN
  the assertions were meaningful
  the edge cases were tested
  the code is correct

A test executing code with no assertion counts towards coverage. Coverage measures reach, not verification.

Use it as a discovery tool: low coverage on a critical module is a genuine finding. As a target it produces tests written to hit lines, which are the least useful tests that exist — Goodhart’s law, on schedule.

A workable default for a web application

TYPES              strict, everywhere
                   → the cheapest tier

UNIT               pure logic — pricing,
                   validation, formatting
                   → where the rules are

INTEGRATION        the bulk
                   → components with real
                     rendering, API routes
                     with a real database

E2E                a handful of critical
                   journeys
                   → browse → basket →
                     checkout → confirm
                   → run on every deploy

VISUAL             the design system
                   — Visual Regression Testing

MANUAL             exploratory, on anything
                   genuinely new

See: Visual Regression Testing

The critical-journey E2E suite is the one to have first if you have nothing. Half a dozen tests covering the money path catch the failures that matter most, and they justify the rest of the pipeline — End-to-End Testing.

What not to test

Stated explicitly, because it’s rarely written down:

  • The framework. It has its own tests
  • Third-party libraries
  • Trivial getters and setters
  • Implementation details — testing internals makes refactoring break tests, which teaches people to avoid refactoring — Unit Testing
  • Anything you’d delete rather than fix