Field vs Lab Data
Date: 2026-08-16
Field data is what your customers experienced; lab data is what happened in a controlled simulation. They disagree by design — field tells you whether you have a problem, lab tells you why, and using either for the other’s job is the field’s most common mistake.
What they are
- Field data (real user monitoring, RUM) — measurements collected from actual visits, on their devices, connections and locations
- Lab data (synthetic) — measurements from a scripted run in a controlled environment, with a fixed device profile and throttled network
| Field | Lab | |
|---|---|---|
| Reflects real users | Yes | No |
| Reproducible | No | Yes |
| Detailed diagnostics | Limited | Extensive |
| Available for a new page | No — needs traffic | Yes, immediately |
| Can measure INP meaningfully | Yes | Barely |
| Good for regression testing | Poor — noisy, delayed | Excellent |
| What Google grades you on | Yes | No |
Why they disagree
Not a fault in either — they measure different populations.
- Device mix. A lab run uses one profile. Your field data is every phone your customers own, including five-year-old Androids
- Network. Throttling approximates a connection; it doesn’t reproduce packet loss, congestion or a train tunnel
- Cache state. Lab runs are usually cold. A large share of real visits are warm — HTTP Caching
- Behaviour. Real users scroll, tap and navigate away. A script doesn’t, which is why INP barely exists in a lab
- Third parties. Consent state, ad blockers and extensions change what actually loads — Third-Party Scripts
- Geography. Distance to origin is fixed by physics and varies by visitor
In plain terms: a Lighthouse score is a single run on a made-up phone. Your field data is thousands of runs on real ones. When they disagree, the field is right about what’s happening and the lab is right about what’s causing it.
Using each properly
Field data for:
- Deciding whether there’s a problem at all
- Prioritising — which template, which device, which country is worst
- Tracking whether a fix actually helped real people
- Anything you report to stakeholders
Lab data for:
- Diagnosing a known problem in detail — waterfalls, flame charts, phase breakdowns
- Testing a fix before deploying
- Regression testing in CI, where reproducibility is the whole point
- Pages with no traffic yet
The workflow that follows: field to find it, lab to fix it, field to confirm. Skipping the last step is how teams end up with a great Lighthouse score and unchanged customer experience.
The 28-day lag
Chrome’s public field data (CrUX) is a rolling 28-day aggregate at the 75th percentile. So:
- A fix deployed today doesn’t fully appear for four weeks
- A regression is equally slow to show
- Comparing two adjacent weeks compares two heavily overlapping windows
Your own Real User Monitoring has no such lag and can be segmented by anything you collect, which is why it’s worth running alongside the public sources rather than instead of them.
The traps
- Optimising the Lighthouse score. It’s a weighted composite of lab metrics. Improving it is not the same as improving the experience, and some things that raise it — deferring everything indiscriminately — hurt real users
- One lab run as evidence. Lab results vary run to run. Take a median of five, minimum
- Testing on your own machine and connection. Throttle CPU to 4–6× and network to a slow 4G profile, or you’re measuring a device none of your customers has
- Comparing lab to field thresholds. The 2.5s LCP threshold is a field p75. A lab run showing 2.4s doesn’t mean you pass
- Averaging field data. Performance distributions are heavily skewed; the mean is meaningless — Percentiles in Performance
What to run
- Field: CrUX via PageSpeed Insights and Search Console for the graded view; your own RUM for the actionable one
- Lab: Lighthouse for a quick diagnostic, DevTools Performance panel for real investigation, WebPageTest for filmstrips and connection-level detail
- CI: a lab tool with a fixed configuration and a threshold that fails the build — Performance Budgets