Tags: web-dev concept

Pipeline Design

Date: 2026-08-17


Arranging CI stages so feedback arrives fast and the whole thing stays under the attention span it’s competing with. Total runtime is the number people quote; time-to-first-failure is the one that determines whether the pipeline gets used.


Pipeline design is the arrangement of stages, their order, what runs in parallel, what’s cached, and what blocks a merge.

Two numbers

TOTAL RUNTIME          how long until
                       everything is green

TIME TO FIRST FAILURE  how long until you
                       learn something is
                       wrong
                       ← the one that shapes
                         behaviour

Optimise the second first. A pipeline where a lint error surfaces in 20 seconds is far more usable than one where everything finishes in six minutes but nothing reports until the end.

Order by speed

0:00   install (cached)
0:20   ├─ lint            ┐
0:25   ├─ format check    │ parallel
0:40   ├─ typecheck       │
0:45   └─ unit tests      ┘
1:30   build
3:00   ├─ integration     ┐ parallel
5:00   └─ e2e (critical)  ┘

Cheap and fast checks first, so the common mistakes fail immediately. Everything independent runs in parallel, so the pipeline’s length is its longest path rather than its sum.

Fail fast, but not too fast

FAIL-FAST ON
  lint, typecheck — the fix is obvious
  and immediate

DON'T FAIL-FAST ON
  test suites
  → you want to know ALL the failures,
    not just the first
  → cancelling the rest means fixing
    one and rediscovering the next

A test job should run to completion and report every failure. Cancelling parallel siblings on the first failure turns one round trip into four.

Caching

The largest single lever on runtime — Build Caching:

DEPENDENCIES     keyed on the lockfile hash
BUILD OUTPUT     keyed on source + config
TEST RESULTS     skip unaffected packages
DOCKER LAYERS    order the Dockerfile
                 correctly

Dependency caching is the cheapest and most often missing:

- uses: actions/setup-node@v4
  with: { node-version: 24, cache: npm }
- run: npm ci

[CHECK: the current major of any action you pin, and the current Node LTS. Both move, and a pinned action major eventually stops receiving updates.]

Run only what’s affected

On a monorepo, the biggest win available:

CHANGED   packages/checkout

RUN       checkout's tests
          + anything depending on it

SKIP      everything else

Turns a fifteen-minute pipeline into a two-minute one on most pull requests — Monorepos.

Blocking versus reporting

BLOCKING — must be fast and deterministic
  lint, format, typecheck
  unit + integration tests
  build
  critical E2E journeys
  secret scanning

REPORTING — informs, doesn't stop
  bundle size trend
  coverage trend
  full accessibility scan
  dependency audit
  visual diffs
  performance budgets, as a warning
    first

Putting a slow or flaky check in the blocking set is how pipelines lose trust. Anything non-deterministic belongs in the reporting set until it’s reliable — Flaky Tests.

Artefacts and environments

BUILD ONCE
   │
   ├─→ test against THAT artefact
   ├─→ deploy to preview
   └─→ deploy to production

Never rebuild between environments. Building separately for staging and production means the thing you tested isn’t the thing you shipped, and the difference is exactly where environment-specific bugs live — Continuous Deployment.

Making failures diagnosable

The part that gets neglected until an incident:

  • Upload artefacts on failure — screenshots, traces, logs, coverage
  • Playwright traces are the single most useful E2E artefact: a full timeline with DOM snapshots
  • Name jobs for what they check, not for the tool. “Checkout E2E” beats “job-3”
  • Annotate the failure in the pull request, so nobody opens logs to find a lint error

The realistic target

BLOCKING PIPELINE        under 10 minutes
TIME TO FIRST FAILURE    under 1 minute
FLAKINESS                near zero

Under ten minutes is the threshold where people wait for it. Above twenty, they merge before it finishes and the gate stops functioning — Continuous Integration.

When it exceeds the budget, the order to attack is: cache, parallelise, run only what’s affected, then move the slowest thing to the reporting set.