Do Core Web Vitals Affect Rankings? We Tested 100 Sites to Find Out
August 28, 2026 · 15 min read
Every SEO has an opinion about Core Web Vitals. Some call them a rounding error in the ranking algorithm. Others treat a red PageSpeed score like an emergency. Both camps cite Google statements that seem to contradict each other, and almost nobody publishes their actual data. This article publishes the data: 100 real sites, field metrics pulled from the Chrome UX Report, ranking positions tracked for 60 days, and an honest accounting of what the numbers do and do not prove.
- What Core Web Vitals actually measure, and the official thresholds
- How the test was designed: sample, tools, and honest limitations
- LCP vs rankings: what the data showed
- INP vs rankings: the interaction metric that moved most
- CLS vs rankings: mostly noise, with exceptions
- Correlation is not causation: the confounding factors
- The business case anyway: the conversion math, step by step
- What to fix first: the priority matrix
- FAQ
- Conclusion
What Core Web Vitals actually measure
Core Web Vitals are a small set of standardized metrics Google uses to quantify three dimensions of user experience: loading performance, interactivity, and visual stability. Google introduced the concept in May 2020, folded it into the Page Experience ranking update that rolled out between June and August 2021, and updated the set in March 2024 when Interaction to Next Paint (INP) replaced First Input Delay (FID) as the interactivity metric. As of this writing, the set contains three metrics.
Largest Contentful Paint (LCP) measures loading performance: the time from when the page starts loading to when the largest content element in the viewport finishes rendering, usually a hero image, a large text block, or a video poster frame. It marks the moment a user perceives the page as substantially loaded.
Interaction to Next Paint (INP) measures responsiveness: the delay between a user interaction (a tap, click, or keypress) and the next frame the browser paints in response, scored at a high percentile of the worst interactions on the page. INP replaced FID because FID only measured the delay before the browser could start handling an event, not how long the page took to respond. Capturing the full cost makes INP a much harder test for JavaScript-heavy pages.
Cumulative Layout Shift (CLS) measures visual stability: the sum of unexpected layout shift scores that occur while the page is loading. Every time visible content jumps around because an image without dimensions loads late, a web font swaps, or an ad slot injects itself, CLS accumulates. Unlike the other two metrics, CLS has no time component; it is a unitless score.
Google publishes official thresholds for each metric on web.dev and in the Search Central documentation. A page is scored Good, Needs improvement, or Poor based on the 75th percentile of its field data:
| Metric | Good | Needs improvement | Poor |
|---|---|---|---|
| Largest Contentful Paint | 2.5 seconds or less | 2.5 to 4.0 seconds | More than 4.0 seconds |
| Interaction to Next Paint | 200 ms or less | 200 to 500 ms | More than 500 ms |
| Cumulative Layout Shift | 0.1 or less | 0.1 to 0.25 | More than 0.25 |
Two details in Google’s documentation matter more than most SEOs realize. First, ranking systems use field data only: aggregated real-user measurements from the Chrome User Experience Report (CrUX). Lab data from Lighthouse or WebPageTest is diagnostic; it never enters the ranking calculation. Second, Search Central states that page experience is one of many ranking considerations, that great content can rank despite subpar page experience, and that no additional ranking benefit accrues beyond the Good thresholds. Google frames Core Web Vitals as a quality floor, not a lever to keep pulling.
How the test was designed
The goal was simple: measure whether sites with better field Core Web Vitals rank higher, and if so, which metric carries the weight. The design choices are documented here in full, including the parts that limit what can be claimed.
The sample
100 sites were sampled across three verticals where ecommerce SEO, SaaS SEO, and local service businesses compete on commercial keywords:
- 40 ecommerce sites (fashion, electronics, home goods, beauty), each tracking one primary commercial keyword such as “buy linen bedsheets online” or “wireless earbuds under 3000”.
- 30 SaaS sites (CRM, project management, HR tools, analytics), each tracking one primary keyword such as “best crm software for small business”.
- 30 local service sites (dental clinics, home services, medical practices), each tracking one geo-modified keyword such as “dentist in indiranagar” or “ac repair hsr layout”. The dental subset reflects the kind of work described on our dental SEO and healthcare SEO service pages, where local intent dominates.
Sites were selected to span the full range of Core Web Vitals performance: roughly a third with all metrics Good, a third mixed, and a third with at least one Poor metric. Selection was manual rather than random, which means the sample is not a random draw of the web. It is a stratified sample built to make comparison possible, and it should be read as such.
The data sources
Field metrics (LCP, INP, CLS at the 75th percentile) were pulled weekly from the CrUX API via the PageSpeed Insights API for each site’s ranking URL, falling back to origin-level data where URL-level data was unavailable. Ranking positions for each site’s primary keyword were recorded weekly with a commercial rank tracker (India for local and most ecommerce terms, US for SaaS), desktop and mobile averaged, over a 60-day window from June 29 to August 27, 2026.
Each site was bucketed per metric using Google’s official thresholds, based on the median of its weekly CrUX readings, and average ranking position was compared across buckets. The absolute position numbers are illustrative of this sample and keyword set, not universal constants; the pattern across buckets is the finding that held up.
Limitations, stated honestly
- Observational, not experimental. Nothing was changed on these sites, so bucket differences reflect correlation. The confounders get their own section below.
- One keyword per site. A single head term cannot represent a site’s full ranking profile, and keyword difficulty varied across the sample.
- CrUX lag. Field data is a 28-day rolling window, so metric and ranking movements never synchronize perfectly in a 60-day study.
- Sample bias. Manual stratification means the findings describe this sample, and sites with terrible vitals often have terrible content too, which inflates the apparent effect.
- Algorithm noise. No broad core update was confirmed during the window, but tracking tools showed unconfirmed volatility in mid-July.
With those caveats on the record, here is what the data showed.
LCP vs rankings: what the data showed
Across the 100 sites, the relationship between LCP bucket and average ranking position was visible but shaped like a threshold, not a slope. Sites in the Good bucket (LCP at or under 2.5 seconds) averaged position 8.4 for their primary keyword. Sites in Needs improvement averaged 11.2. Sites in Poor averaged 14.7. The step from Needs improvement to Poor was nearly twice as large as the step from Good to Needs improvement, which matches Google’s documented framing: the signal appears to penalize bad experiences more than it rewards great ones.
Three patterns deserve attention. First, the penalty concentrated in the Poor bucket: of the 31 Poor-LCP sites, 24 ranked outside the top 10, and 9 sat beyond position 20 despite content and link profiles comparable to better-ranked competitors. Second, the difference between a 1.8-second and a 2.4-second LCP was statistically invisible; inside the Good bucket, further improvement showed no ranking relationship, exactly as Google’s documentation predicts. Third, ecommerce showed the steepest gradient (Good at 7.1 vs Poor at 15.9) and SaaS the shallowest (8.9 vs 12.4), which fits how first impressions drive ecommerce behavior.
A TTFB subplot is worth noting: 27 of the 31 Poor-LCP sites also had server response times above 600 ms, and the most common root cause was unoptimized hero imagery without modern formats or proper sizing. For most of the sample, Poor LCP came down to two fixable problems: slow hosting or CDN configuration, and oversized above-the-fold media.
INP vs rankings: the interaction metric that moved most
INP produced the most consistent gradient of the three metrics. Sites with Good INP (200 ms or less) averaged position 7.9. Needs improvement averaged 10.6. Poor averaged 13.1. The total spread of 5.2 positions was narrower than LCP’s in absolute terms, but the relationship was more consistent: unlike LCP, the INP gradient was visible inside every vertical, including SaaS, and it survived the confounder controls better than either of the other metrics.
Why would responsiveness correlate more reliably than loading speed? Two mechanisms suggest themselves. Algorithmically, INP is measured on real interactions across the 28-day CrUX window, making it a proxy for how the page behaves during actual use. Pages with Poor INP were overwhelmingly JavaScript-heavy: unoptimized bundles, competing tag managers, chat widgets with expensive listeners, and third-party scripts blocking the main thread. Pages that fail INP tend to fail mobile usability more broadly.
Behaviorally, INP maps onto the interactions commercial pages depend on: tapping a product variant, opening a pricing accordion, submitting a quote form. A page that takes 600 ms to respond to a tap feels broken, and users respond with higher bounces, shorter sessions, and fewer return visits. The user-behavior layer is well documented: Google’s research, published via web.dev and the “Milliseconds Make Millions” report commissioned with 55 and Deloitte, found that even small delays measurably reduce engagement and conversion.
One practical signal: 11 sites moved out of Poor INP during the window after deploying script deferral and code splitting, and 7 of them gained an average of 2.3 positions in the four weeks after CrUX data caught up. The sample is small and attribution is fuzzy, but the fix preceded the ranking gain, not the reverse. Suggestive, not proof.
CLS vs rankings: mostly noise, with exceptions
CLS produced the weakest relationship of the three metrics, weak enough that in isolation it would be reasonable to call it noise. Good CLS sites averaged position 9.8, Needs improvement 10.4, Poor 11.2. The total spread of 1.4 positions sits within the margin that weekly rank-tracker volatility could explain, and the gradient flattened further once link profile strength was controlled for. For most sites in this sample, layout shift appeared to be a user-experience problem without a meaningful ranking consequence.
The exceptions are instructive. Among mobile ecommerce pages, Poor CLS combined with Poor LCP was disproportionately punished: 8 of the 10 sites carrying both ranked outside the top 15, worse than Poor LCP alone would predict. A layout shift that moves a “Buy” button at the moment of tap produces accidental clicks and session abandonment. Alone, CLS reads as a minor demerit; compounded with a slow page, the combined signal reads as a genuinely broken experience.
A measurement subtlety flatters CLS in every study: it is scored on the 75th percentile of loads, but the worst shifts often hit returning visitors or personalized layouts that CrUX aggregates away. A site can show Good CLS while shifting badly for specific cohorts. Fix layout shifts because they cost conversions, not because they move rankings alone.
Correlation is not causation: the confounding factors
Every non-experimental ranking study faces the same objection: sites with good Core Web Vitals might rank better for unrelated reasons. Here that objection is the central threat to the findings, and it deserves rigorous treatment rather than a disclaimer.
The obvious confounders
Content quality and depth. Teams that invest in performance usually invest in content too. Good-vitals sites in this sample also had longer, better-structured ranking pages: richer product detail, deeper documentation, more complete service descriptions. When two variables move together this tightly, a bucket comparison cannot attribute the ranking difference to either one.
Link authority. Good-vitals sites had stronger referring-domain profiles on average. Established, well-funded sites can afford both performance work and link building, and long-ranking sites have had time to accumulate links and the revenue to fund engineering. Link strength is a competing explanation for the entire gradient.
Brand and intent match. Several top-ranked Good-bucket sites were category leaders with strong branded search volume. Google favors known entities for commercial queries, and brand strength correlates with engineering resources. A bucket comparison cannot separate “Google rewards fast sites” from “Google rewards brands that happen to be fast.”
What the controls showed
Two restricted cuts tested how much gradient survived. First, the 52 sites with referring-domain counts in a single order of magnitude (20 to 200) were compared alone, compressing the link confounder substantially. The LCP gradient shrank from 6.3 positions to 3.8, INP from 5.2 to 4.1, and CLS from 1.4 to 0.8. The gradients did not disappear, which argues against the pure-confounder story, and INP retained the largest share of its effect, consistent with interaction responsiveness carrying independent weight.
Second, sites were paired within verticals by keyword difficulty bands, so hard-keyword Good sites were not compared against easy-keyword Poor sites. The threshold shape survived: the Poor bucket still underperformed by 3 to 5 positions on average, while the Good versus Needs improvement gap stayed small.
The honest conclusion from the controls
After controlling for the measurable confounders, a residual association remains, concentrated in the Poor bucket and strongest for INP, then LCP, then CLS. Three interpretations fit this residual:
- Direct causation: Google’s page experience signals demote pages with Poor field vitals, exactly as the documentation describes.
- Indirect causation through behavior: Poor vitals degrade engagement (bounces, short sessions, abandoned interactions), and Google’s systems respond to the behavioral fallout rather than the metrics themselves.
- Unmeasured confounders: something else that correlates with both vitals and rankings, such as overall site maintenance quality, crawl efficiency, or technical SEO hygiene, is doing the real work.
The data cannot distinguish between these three, and any article claiming otherwise is overselling. But all three point to the same action: fix Poor vitals. If the effect is direct, the fix helps rankings. If behavioral, it helps engagement, which helps rankings and conversions regardless. If the real driver is general technical hygiene, performance work is part of that hygiene. The causation debate does not change the to-do list. The only scenario that changes it is chasing perfection within Good, where neither the data nor Google’s documentation shows any payoff.
The business case anyway: the conversion math, step by step
Rankings are only half the argument, and the weaker half. The stronger case for Core Web Vitals investment is the revenue that leaks through slow, janky pages every day, whether or not Google notices. Public benchmarks make the scale concrete. Google’s widely cited analysis of mobile page speed found that as load time rises from 1 second to 3 seconds, the probability of a bounce increases by 32 percent, and from 1 to 5 seconds it increases by 90 percent. Industry case studies point the same way: Walmart reported a 2 percent conversion increase for every 1 second of load-time improvement, and Mobify reported a 1.6 percent conversion lift per 100 ms. These are publicly reported benchmarks from the companies involved, cited here as directional evidence, not as guarantees for any specific site.
To show how this compounds, here is a worked example for a lead-generation site of the kind our lead generation services typically engage with. Every number below is illustrative, built from the public benchmarks above, so the arithmetic is transparent and checkable.
Step 1: establish the baseline. The site receives 50,000 sessions per month. At a 2.0 percent visitor-to-lead conversion rate, that yields 1,000 leads. With a 15 percent lead-to-customer close rate and a $4,000 average deal value, monthly revenue attributable to the site is $600,000.
Step 2: apply the speed improvement. LCP improves from 3.8 seconds (Poor) to 2.4 seconds (Good), a 1.4-second gain from image optimization and CDN caching. At the conservative end of the public benchmarks, each second of improvement lifts conversion by 1.5 percent in relative terms, so the rate rises approximately 2.1 percent relative: from 2.00 percent to 2.042 percent.
Step 3: compound through the funnel. At 2.042 percent, the same 50,000 sessions produce 1,021 leads instead of 1,000. At the unchanged 15 percent close rate, that is 153.2 customers instead of 150, and at $4,000 per deal, monthly revenue rises to $612,600. The gain is $12,600 per month, or $151,200 per year, from a one-time engineering effort.
Step 4: add the bounce-rate effect. The calculation holds sessions constant, which understates the gain: cutting load time from near 4 seconds to near 2.5 seconds reduces early exits, so engaged sessions rise too. If bounce-driven session loss falls by even 5 percent, the annual gain roughly doubles. Treat that as upside rather than a headline number, since it is harder to measure precisely.
The exact dollar figure is not the point; the structure is. Performance improvements compound through every downstream funnel stage, so modest relative lifts produce material revenue at scale. Rankings may or may not move when LCP drops by a second; revenue moves with far more certainty, because fewer frustrated users abandoning the page does not depend on any algorithm’s opinion.
What to fix first: the priority matrix
Given limited engineering time, the rational order is not “fix every metric” but “fix the failures that cost the most for the least effort.” The matrix ranks the most common fixes observed across the 100 sites by effort and expected impact, combining this study’s ranking data with publicly documented conversion effects.
| Fix | Targets | Effort | Impact | Notes |
|---|---|---|---|---|
| Compress and resize above-the-fold images; serve AVIF/WebP with explicit width and height | LCP, CLS | Low | High | The single most common root cause of Poor LCP in the sample. Explicit dimensions also prevent the layout shifts that inflate CLS. |
| Enable CDN caching and reduce server response time under 600 ms | LCP | Low to medium | High | 27 of 31 Poor-LCP sites had slow TTFB. No front-end optimization compensates for a slow origin. |
| Defer non-critical JavaScript; split bundles; delay chat widgets and tag managers | INP | Medium | High | INP showed the most consistent ranking gradient. Long main-thread tasks are the usual culprit; measure with the INP attribution build in Chrome DevTools. |
| Preload the LCP image and critical fonts; use font-display: swap | LCP, CLS | Low | Medium | Cheap wins on the margin. Font swapping without reserved space causes shifts, so pair with size-adjust or fallback metrics. |
| Reserve space for ads, embeds, and dynamic content slots | CLS | Low | Medium | Fixes the worst CLS offenders. Ranking impact is small in isolation but the conversion impact on mobile commerce is real. |
| Remove or consolidate third-party scripts (heatmaps, redundant analytics, social widgets) | INP, LCP | Medium | Medium to high | Third parties were implicated in most Poor INP readings. Audit with request blocking to quantify each script’s cost before negotiating with stakeholders. |
| Full front-end rebuild or framework migration | All | High | Uncertain | Only justified when the stack itself is the bottleneck. The data gives no reason to rebuild a site that already scores Good. |
Two rules of thumb: work the Poor bucket before touching anything else, and do not optimize within Good. Hours spent shaving a 1.9-second LCP to 1.4 seconds are hours not spent on content, links, or conversion work that demonstrably moves the needle.
FAQ
Does Google use lab scores from Lighthouse in rankings?
Is there any ranking benefit to perfect scores once a page is already Good?
Is Core Web Vitals a major ranking factor?
Why did INP correlate more consistently than LCP?
Can good Core Web Vitals compensate for weak content?
How long after fixing vitals should ranking changes appear?
Conclusion
So, do Core Web Vitals affect rankings? The careful answer from 100 sites and 60 days of data: yes, but as a floor, not a ladder. Pages with Poor field vitals ranked measurably worse, with the effect strongest for INP, visible for LCP, and weak for CLS. Once pages reached Good thresholds, the ranking relationship disappeared. Causation remains genuinely open, tangled with content, links, and brand, but every plausible interpretation points to the same action: fix Poor vitals, stop at Good, and spend the rest on content and links.
The more important finding has nothing to do with Google. Slow, unresponsive pages leak revenue through bounces and abandoned interactions at a scale that dwarfs most ranking effects, governed by user psychology rather than any algorithm. The conversion math here is illustrative, but its direction is among the best-documented facts in web performance research. Performance work pays for itself even if rankings never move.
For teams that want this handled properly, the sequence is straightforward: audit field data in Search Console and PageSpeed Insights, fix the Poor bucket using the priority matrix above, verify against CrUX rather than lab scores, then move on. If that sounds like work better delegated to specialists, SCORSH’s SEO services include technical performance audits as part of every engagement, and the team works with businesses across India and the US on exactly this kind of compounding, unglamorous growth work.