Tags: web-dev concept

Synthetic Monitoring

Date: 2026-08-17


Repeatable lab runs on a schedule, from a fixed environment. It can’t tell you what customers experienced — that’s Real User Monitoring (RUM) — but it’s the only way to get a clean before-and-after on one change, because it holds everything else constant.


Synthetic monitoring is running scripted page loads or journeys on a schedule from controlled devices, locations and networks, and recording the results over time.

The job it does that RUM can’t

RUM                                  SYNTHETIC

real users, real conditions          one scripted run, fixed conditions
every variable moves at once         everything held constant except
                                       your change
answers: do we have a problem,       answers: did THIS change help,
  and for whom                         and why
lagging — needs traffic to           immediate — run it on a branch
  accumulate                           before merging
can't test what doesn't exist yet    tests a preview deploy
no waterfall, no trace               full waterfall, filmstrip, trace

Synthetic is a controlled experiment on your own site, with n=1 and every confounder pinned. That’s exactly what you want for attribution of a change, and exactly what you can’t get from field data where traffic mix, device mix and network conditions shift constantly — Field vs Lab Data.

The variance problem

The thing that makes synthetic monitoring unreliable when done naively: the same page, run twice, gives different numbers.

ten consecutive runs of an unchanged page
LCP — largest contentful paint, in seconds

2.41  2.38  2.67  2.44  2.39  2.91  2.42  2.40  2.55  2.43

median 2.42    min 2.38    max 2.91    spread 0.53s

a "regression" of 0.3s is inside the noise.
a single run showing 2.91 means nothing.

Sources: CPU contention on shared CI runners, network variability, third-party scripts responding at different speeds, and A/B tests assigning the run to different variants.

The mitigations, in order of effect:

  • Median of at least 5 runs, ideally 9. Never a single run, ever
  • Compare against a baseline run in the same session, not against a stored historical number. Run the base branch and the change branch back to back on the same machine — this cancels most environmental drift
  • Pin what you can — CPU throttling multiplier, network profile, viewport, device emulation
  • Stub or block third parties for the comparison run. You’re measuring your change, and a slow tag server adds a second of unattributable noise
  • Force a single test variant, or your runs are sampling from two different pages — Experiment QA
  • Dedicated hardware over shared CI runners where the budget allows. Shared runners are the single largest variance source

Report the median and the spread. A result of “2.42s, spread 0.53” is honest; “2.42s” implies a precision that doesn’t exist — Percentiles in Performance, Communicating Uncertainty.

What to run, and how often

ON EVERY PULL REQUEST      the key templates — home, category, product,
                           basket, checkout
                           compared against the base branch
                           ← the highest-value use — Performance Regression Testing

HOURLY / DAILY             a small set of critical pages in production
                           trend detection, third-party drift,
                           "did something change that wasn't us"

ON A SCHEDULE, MULTI-      from the geographies you sell to
GEOGRAPHY                  catches CDN and DNS problems invisible
                           from one location

AFTER EVERY DEPLOY         a smoke run, to catch the obvious

The production schedule catches what CI can’t: things that changed without a deploy. A third-party script getting heavier, a CDN configuration drift, an origin slowdown. These produce no commit and no alert until a customer complains.

What it’s bad at

  • Representing your users. One device profile, one connection, one location, cache cold. Real traffic is a distribution and synthetic is a point — never quote a synthetic number as “our LCP”
  • Anything requiring real state. A logged-in journey with a real basket needs scripted authentication and test data that stays valid, which is where synthetic suites rot — Test Data
  • Long sessions. Every synthetic run is a fresh load, so it structurally cannot see degradation over time — Memory and Long Sessions
  • Personalisation and geography. The run gets one variant from one place
  • Rare conditions. The slow 3G connection that 4% of your customers have won’t be represented unless you configure a profile for it, and then it’s not representative of the other 96%

Using both properly

FIELD (RUM) says          →  "p75 LCP on product pages is 3.4s,
                              worst on Android in the North West"

SYNTHETIC then says       →  "here's the waterfall: the hero image
                              starts at 1.9s because it's discovered
                              after the CSS"

fix, verify in SYNTHETIC  →  "median LCP 2.1s, spread 0.2, vs 2.9
  against a baseline run      on the base branch"

confirm in FIELD          →  "p75 moved to 2.6s over the next 28 days"

That loop is the whole practice. Field to find and to confirm; synthetic to diagnose and to verify. Using either alone produces the two standard failures — optimising things no customer experiences, or making changes you can’t attribute.

Where it interacts