Tags: web-dev concept

Performance Regression Testing

Date: 2026-08-17


Catching a slowdown in the pull request rather than in the field a fortnight later. The hard part isn’t the threshold — that’s Performance Budgets — it’s that the measurement is noisy enough to produce false failures, and a check that cries wolf gets disabled within a month.


Performance regression testing is measuring performance automatically on every change, before merge, and failing the change when it breaches a threshold.

Why the field is too late

merge  →  deploy  →  field data accumulates  →  p75 moves  →  someone notices
  ↑                        28-day rolling window
  │                                    ↓
  │                            ~3–6 weeks later
  │
  └── by now: 40 more commits, and nobody can say which one did it

Core Web Vitals field data uses a 28-day window — those being LCP (largest contentful paint), INP (interaction to next paint) and CLS (cumulative layout shift), so a regression is diluted for weeks before it’s visible and then can’t be attributed to a commit. The whole argument for CI measurement is attribution, not speed of detection — Field vs Lab Data.

Two things to check, and one is much cheaper

STATIC ANALYSIS — bundle size          RUNTIME — Lighthouse / a real run

fast (seconds), deterministic          slow (minutes), noisy
no variance at all                     needs medians and a baseline
catches: payload growth,               catches: everything else —
  a new dependency, a chunk              blocking resources, layout
  moving to the wrong route              shift, slow interactions

  ← run on EVERY PR. no excuse           ← run on PRs touching the
                                           front end, or nightly

Bundle size checks are the ones to have first. They’re free, exact, and catch the most common regression — someone imported a library. A build that fails at +40 KB is a five-minute conversation; the same 40 KB found in a quarterly audit is archaeology — Bundle Analysis.

The comparison that survives noise

Absolute thresholds fail in CI because the runner’s speed varies day to day. Compare against the base branch in the same run.

✗  ABSOLUTE
     fail if LCP > 2.5s
     → passes on a fast runner, fails on a busy one,
       for identical code

✓  RELATIVE
     run base branch:  median LCP 2.38s  (9 runs)
     run this branch:  median LCP 2.44s  (9 runs)
     delta +0.06s, inside the run-to-run spread → pass

     fail only if the delta exceeds the measured noise floor

Establish the noise floor empirically. Run the same commit against itself ten times and record the spread; anything smaller than that is not a signal. On typical shared CI this is often ±0.2–0.3s for LCP, which means small regressions are genuinely undetectable there — and saying so is better than pretending otherwise.

The full recipe, since getting this wrong is why these checks get abandoned:

  • Median of 5–9 runs, never one
  • Both branches in the same job, on the same machine, back to back
  • Third parties blocked or stubbed — you’re testing your change
  • A/B tests forced to one variant
  • Fixed CPU throttle and network profile
  • Cache state explicit — cold load for first-visit metrics, warm for repeat

What to gate, and how hard

HARD FAIL — blocks the merge          WARN — comments, doesn't block

bundle size over budget               LCP delta beyond the noise floor
a new render-blocking resource        a new third-party origin appears
LCP element became lazy-loaded         CLS delta
a dependency added to the             INP delta
  initial chunk
                                      ← too noisy to block on, still
← deterministic. no false positives     worth surfacing in the PR

Only block on deterministic checks. A hard gate on a noisy runtime metric produces failures on unrelated pull requests, which produces a culture of re-running until green, which is worse than no check at all because it also teaches people to ignore the ones that matter.

Post the result as a PR comment either way. A table showing base versus branch, with deltas, puts the information where the decision is being made — and most regressions get fixed on sight when someone can see they caused them, without any gate at all.

Making it stick

  • Start with bundle size only. It works immediately, never false-fails, and buys credibility for adding the noisy checks later
  • Test the pages that matter, not the homepage. Product, category and checkout are where the revenue is
  • Give it an owner. An unowned check that starts failing gets disabled — Performance Culture
  • Review the budget quarterly. A threshold set at last year’s baseline either strangles the team or has been ratcheted up so often it means nothing — Performance Budgets
  • Track the check’s own reliability. If a fifth of failures are false, fix the harness before adding anything to it
  • Pair it with a production schedule, because CI cannot catch what changed without a deploy — a third party getting heavier, a CDN config drift — Synthetic Monitoring

The limits worth stating

CI measurement catches what you scripted, on the templates you chose, in a lab environment. It will not catch a regression only visible on real devices, only on one geography, only for logged-in users with 40 items in a basket, or only after twenty minutes of use — Memory and Long Sessions.

So the field remains the source of truth and CI is the attribution mechanism. A regression that reaches production is a gap in the suite worth closing — add the case rather than concluding the suite doesn’t work.

Where it interacts