A green 100 is satisfying and it is worth having. It is also a claim about one specific thing: a single scripted cold load of one URL, on one simulated device, over one simulated network, with no user in it.
Real visitors are a distribution of devices, connections, locations, entry points and behaviours. When a site scores 100 in the lab and still fails Core Web Vitals in the field, it is almost never because Lighthouse lied. It is because the question it answered was narrower than the question you cared about.
Here is what that run cannot see.
1. Nobody interacted
Lighthouse has no INP score. It cannot have one, because INP measures the latency of real interactions and a scripted load does not click anything.
What it reports instead is Total Blocking Time — main-thread blocking during load. TBT can help identify load-time work that may also affect responsiveness, but it is not an INP measurement. A 0 ms TBT tells you the page was quiet while it loaded; it says nothing about the filter panel, menu, carousel or modal if the run never opened them.
2. The simulated device is one point on a curve
The mobile preset simulates a mid-range Android on a throttled connection. Your audience contains phones considerably slower than that, on networks considerably worse, and Core Web Vitals is assessed at the 75th percentile — it is explicitly designed to notice the slowest quarter of your visits. A single simulated device cannot represent a distribution, and the point of the percentile is the tail.
3. Only one URL was tested
Almost everyone runs Lighthouse on the homepage. The homepage is usually the most optimised page on the site, and often the least visited.
Field data in Search Console groups URLs and reports on the group. Your product pages, your search results, your paginated listings and your blog templates are in there, and they may share none of the homepage’s careful work.
4. Third parties behave differently for real people
A lab run may be geo-located somewhere your tag manager loads a smaller configuration, may not fire consent-gated scripts because there is no consent dialogue to accept, may be excluded from A/B tests, and will not trigger anything keyed to a returning-visitor cookie.
Real users accept the banner, get bucketed into experiments, and receive the chat widget, the heatmap recorder and the retargeting pixel. That is a different page from the one that scored 100.
5. Layout shift accumulates over a visit
CLS is accumulated from unexpected shifts during a page visit, using session windows rather than one unrestricted total. A run that stops after a few seconds and never scrolls may miss an ad slot that resizes below the fold, a lazy image without dimensions, or a sticky banner that appears later.
6. Real navigation is not always a cold load
Lighthouse loads one URL cold, with an empty cache. Real sessions include repeat visitors with a warm cache, back-forward navigations, and — on a client-routed site — soft navigations that skip the document request entirely and are measured differently from what your lab run did.
7. Your origin is fast when only you are asking
A lab run makes a small number of requests under its own conditions. Real traffic arrives in bursts, encounters cache misses, and can find databases under contention. A TTFB measured in one run is not a description of TTFB under all of your traffic.
8. The lab run is now; the field is 28 days
The Chrome UX Report aggregates over a rolling 28-day window. A fix you deployed last week is diluted by three weeks of the old page. This cuts both ways: your green lab score today may be describing a page that field data will not fully reflect for a month, and a regression you shipped yesterday is not visible yet.
What to do about it
Lab testing remains useful, but it answers a narrower question than field data.
- Start from field data. Search Console’s Core Web Vitals report and the CrUX section of PageSpeed Insights tell you whether you have a problem, on which URLs, and for which form factor. That is the scoreboard; Lighthouse is the practice session.
- Measure your own field data. The web-vitals library, in its attribution build, reports which element and which phase caused a bad interaction. Real user monitoring answers questions no lab run can.
- Test the pages people use, not just the homepage. Templates, not URLs.
- Test with interaction, by hand, in the DevTools Performance panel. Open the menu, the filters, the modal. Watch for long tasks.
- Test on a real phone at least once. It can expose slow interactions and loading behaviour that a simulated run does not reproduce.
- Keep the lab score anyway. It is reproducible and it catches regressions before real users meet them, which is exactly what field data cannot do.
I publish lab numbers for my own sites for that reason, and I label them as lab numbers: protectorguardrail.com, Lighthouse 12 on the mobile preset, measured 5 August 2026. They are re-runnable under comparable conditions, which makes them checkable. They do not describe what every visitor experienced; field data is needed for that.
Google explains the distinction in its guide to lab and field data.