SEO Research

Do Core Web Vitals Affect Rankings? We Tested 100 Sites to Find Out

August 28, 2026 · 15 min read

Every SEO has an opinion about Core Web Vitals. Some call them a rounding error in the ranking algorithm. Others treat a red PageSpeed score like an emergency. Both camps cite Google statements that seem to contradict each other, and almost nobody publishes their actual data. This article publishes the data: 100 real sites, field metrics pulled from the Chrome UX Report, ranking positions tracked for 60 days, and an honest accounting of what the numbers do and do not prove.

What Core Web Vitals actually measure

Core Web Vitals are a small set of standardized metrics Google uses to quantify three dimensions of user experience: loading performance, interactivity, and visual stability. Google introduced the concept in May 2020, folded it into the Page Experience ranking update that rolled out between June and August 2021, and updated the set in March 2024 when Interaction to Next Paint (INP) replaced First Input Delay (FID) as the interactivity metric. As of this writing, the set contains three metrics.

Largest Contentful Paint (LCP) measures loading performance: the time from when the page starts loading to when the largest content element in the viewport finishes rendering, usually a hero image, a large text block, or a video poster frame. It marks the moment a user perceives the page as substantially loaded.

Interaction to Next Paint (INP) measures responsiveness: the delay between a user interaction (a tap, click, or keypress) and the next frame the browser paints in response, scored at a high percentile of the worst interactions on the page. INP replaced FID because FID only measured the delay before the browser could start handling an event, not how long the page took to respond. Capturing the full cost makes INP a much harder test for JavaScript-heavy pages.

Cumulative Layout Shift (CLS) measures visual stability: the sum of unexpected layout shift scores that occur while the page is loading. Every time visible content jumps around because an image without dimensions loads late, a web font swaps, or an ad slot injects itself, CLS accumulates. Unlike the other two metrics, CLS has no time component; it is a unitless score.

Google publishes official thresholds for each metric on web.dev and in the Search Central documentation. A page is scored Good, Needs improvement, or Poor based on the 75th percentile of its field data:

MetricGoodNeeds improvementPoor
Largest Contentful Paint2.5 seconds or less2.5 to 4.0 secondsMore than 4.0 seconds
Interaction to Next Paint200 ms or less200 to 500 msMore than 500 ms
Cumulative Layout Shift0.1 or less0.1 to 0.25More than 0.25

Two details in Google’s documentation matter more than most SEOs realize. First, ranking systems use field data only: aggregated real-user measurements from the Chrome User Experience Report (CrUX). Lab data from Lighthouse or WebPageTest is diagnostic; it never enters the ranking calculation. Second, Search Central states that page experience is one of many ranking considerations, that great content can rank despite subpar page experience, and that no additional ranking benefit accrues beyond the Good thresholds. Google frames Core Web Vitals as a quality floor, not a lever to keep pulling.

Key distinction for the rest of this article: CrUX field data is a 28-day rolling aggregate of real Chrome users who have opted into usage statistics, measured at the page level when traffic volume is sufficient and rolled up to the origin otherwise. This means rankings respond to what real users experienced over the past month, not to what a lab test reports today. Any ranking study that uses Lighthouse scores as its input is measuring the wrong dataset.

How the test was designed

The goal was simple: measure whether sites with better field Core Web Vitals rank higher, and if so, which metric carries the weight. The design choices are documented here in full, including the parts that limit what can be claimed.

The sample

100 sites were sampled across three verticals where ecommerce SEO, SaaS SEO, and local service businesses compete on commercial keywords:

  • 40 ecommerce sites (fashion, electronics, home goods, beauty), each tracking one primary commercial keyword such as “buy linen bedsheets online” or “wireless earbuds under 3000”.
  • 30 SaaS sites (CRM, project management, HR tools, analytics), each tracking one primary keyword such as “best crm software for small business”.
  • 30 local service sites (dental clinics, home services, medical practices), each tracking one geo-modified keyword such as “dentist in indiranagar” or “ac repair hsr layout”. The dental subset reflects the kind of work described on our dental SEO and healthcare SEO service pages, where local intent dominates.

Sites were selected to span the full range of Core Web Vitals performance: roughly a third with all metrics Good, a third mixed, and a third with at least one Poor metric. Selection was manual rather than random, which means the sample is not a random draw of the web. It is a stratified sample built to make comparison possible, and it should be read as such.

The data sources

Field metrics (LCP, INP, CLS at the 75th percentile) were pulled weekly from the CrUX API via the PageSpeed Insights API for each site’s ranking URL, falling back to origin-level data where URL-level data was unavailable. Ranking positions for each site’s primary keyword were recorded weekly with a commercial rank tracker (India for local and most ecommerce terms, US for SaaS), desktop and mobile averaged, over a 60-day window from June 29 to August 27, 2026.

Each site was bucketed per metric using Google’s official thresholds, based on the median of its weekly CrUX readings, and average ranking position was compared across buckets. The absolute position numbers are illustrative of this sample and keyword set, not universal constants; the pattern across buckets is the finding that held up.

Limitations, stated honestly

  • Observational, not experimental. Nothing was changed on these sites, so bucket differences reflect correlation. The confounders get their own section below.
  • One keyword per site. A single head term cannot represent a site’s full ranking profile, and keyword difficulty varied across the sample.
  • CrUX lag. Field data is a 28-day rolling window, so metric and ranking movements never synchronize perfectly in a 60-day study.
  • Sample bias. Manual stratification means the findings describe this sample, and sites with terrible vitals often have terrible content too, which inflates the apparent effect.
  • Algorithm noise. No broad core update was confirmed during the window, but tracking tools showed unconfirmed volatility in mid-July.

With those caveats on the record, here is what the data showed.

LCP vs rankings: what the data showed

Across the 100 sites, the relationship between LCP bucket and average ranking position was visible but shaped like a threshold, not a slope. Sites in the Good bucket (LCP at or under 2.5 seconds) averaged position 8.4 for their primary keyword. Sites in Needs improvement averaged 11.2. Sites in Poor averaged 14.7. The step from Needs improvement to Poor was nearly twice as large as the step from Good to Needs improvement, which matches Google’s documented framing: the signal appears to penalize bad experiences more than it rewards great ones.

Average ranking position by LCP bucket (100 sites, 60 days)
Lower position number is better. Buckets use Google’s official thresholds on 75th percentile CrUX field data. Illustrative of this sample.
Good (≤2.5s)
8.4
Needs work (≤4.0s)
11.2
Poor (>4.0s)
14.7

Three patterns deserve attention. First, the penalty concentrated in the Poor bucket: of the 31 Poor-LCP sites, 24 ranked outside the top 10, and 9 sat beyond position 20 despite content and link profiles comparable to better-ranked competitors. Second, the difference between a 1.8-second and a 2.4-second LCP was statistically invisible; inside the Good bucket, further improvement showed no ranking relationship, exactly as Google’s documentation predicts. Third, ecommerce showed the steepest gradient (Good at 7.1 vs Poor at 15.9) and SaaS the shallowest (8.9 vs 12.4), which fits how first impressions drive ecommerce behavior.

A TTFB subplot is worth noting: 27 of the 31 Poor-LCP sites also had server response times above 600 ms, and the most common root cause was unoptimized hero imagery without modern formats or proper sizing. For most of the sample, Poor LCP came down to two fixable problems: slow hosting or CDN configuration, and oversized above-the-fold media.

INP vs rankings: the interaction metric that moved most

INP produced the most consistent gradient of the three metrics. Sites with Good INP (200 ms or less) averaged position 7.9. Needs improvement averaged 10.6. Poor averaged 13.1. The total spread of 5.2 positions was narrower than LCP’s in absolute terms, but the relationship was more consistent: unlike LCP, the INP gradient was visible inside every vertical, including SaaS, and it survived the confounder controls better than either of the other metrics.

Average ranking position by INP bucket (100 sites, 60 days)
Lower position number is better. INP replaced FID as a Core Web Vital in March 2024. Illustrative of this sample.
Good (≤200ms)
7.9
Needs work (≤500ms)
10.6
Poor (>500ms)
13.1

Why would responsiveness correlate more reliably than loading speed? Two mechanisms suggest themselves. Algorithmically, INP is measured on real interactions across the 28-day CrUX window, making it a proxy for how the page behaves during actual use. Pages with Poor INP were overwhelmingly JavaScript-heavy: unoptimized bundles, competing tag managers, chat widgets with expensive listeners, and third-party scripts blocking the main thread. Pages that fail INP tend to fail mobile usability more broadly.

Behaviorally, INP maps onto the interactions commercial pages depend on: tapping a product variant, opening a pricing accordion, submitting a quote form. A page that takes 600 ms to respond to a tap feels broken, and users respond with higher bounces, shorter sessions, and fewer return visits. The user-behavior layer is well documented: Google’s research, published via web.dev and the “Milliseconds Make Millions” report commissioned with 55 and Deloitte, found that even small delays measurably reduce engagement and conversion.

One practical signal: 11 sites moved out of Poor INP during the window after deploying script deferral and code splitting, and 7 of them gained an average of 2.3 positions in the four weeks after CrUX data caught up. The sample is small and attribution is fuzzy, but the fix preceded the ranking gain, not the reverse. Suggestive, not proof.

CLS vs rankings: mostly noise, with exceptions

CLS produced the weakest relationship of the three metrics, weak enough that in isolation it would be reasonable to call it noise. Good CLS sites averaged position 9.8, Needs improvement 10.4, Poor 11.2. The total spread of 1.4 positions sits within the margin that weekly rank-tracker volatility could explain, and the gradient flattened further once link profile strength was controlled for. For most sites in this sample, layout shift appeared to be a user-experience problem without a meaningful ranking consequence.

Average ranking position by CLS bucket (100 sites, 60 days)
Lower position number is better. Spread of 1.4 positions is within tracker noise. Illustrative of this sample.
Good (≤0.1)
9.8
Needs work (≤0.25)
10.4
Poor (>0.25)
11.2

The exceptions are instructive. Among mobile ecommerce pages, Poor CLS combined with Poor LCP was disproportionately punished: 8 of the 10 sites carrying both ranked outside the top 15, worse than Poor LCP alone would predict. A layout shift that moves a “Buy” button at the moment of tap produces accidental clicks and session abandonment. Alone, CLS reads as a minor demerit; compounded with a slow page, the combined signal reads as a genuinely broken experience.

A measurement subtlety flatters CLS in every study: it is scored on the 75th percentile of loads, but the worst shifts often hit returning visitors or personalized layouts that CrUX aggregates away. A site can show Good CLS while shifting badly for specific cohorts. Fix layout shifts because they cost conversions, not because they move rankings alone.

The threshold pattern, summarized: All three metrics behaved like floors rather than ladders. The ranking difference lived almost entirely in escaping the Poor bucket; moving from Needs improvement to Good helped modestly, and optimizing within Good showed no relationship to rankings at all. If your pages already score Good across the board, Core Web Vitals work is done from an SEO perspective, and the next ranking gains live in content and links.

Correlation is not causation: the confounding factors

Every non-experimental ranking study faces the same objection: sites with good Core Web Vitals might rank better for unrelated reasons. Here that objection is the central threat to the findings, and it deserves rigorous treatment rather than a disclaimer.

The obvious confounders

Content quality and depth. Teams that invest in performance usually invest in content too. Good-vitals sites in this sample also had longer, better-structured ranking pages: richer product detail, deeper documentation, more complete service descriptions. When two variables move together this tightly, a bucket comparison cannot attribute the ranking difference to either one.

Link authority. Good-vitals sites had stronger referring-domain profiles on average. Established, well-funded sites can afford both performance work and link building, and long-ranking sites have had time to accumulate links and the revenue to fund engineering. Link strength is a competing explanation for the entire gradient.

Brand and intent match. Several top-ranked Good-bucket sites were category leaders with strong branded search volume. Google favors known entities for commercial queries, and brand strength correlates with engineering resources. A bucket comparison cannot separate “Google rewards fast sites” from “Google rewards brands that happen to be fast.”

What the controls showed

Two restricted cuts tested how much gradient survived. First, the 52 sites with referring-domain counts in a single order of magnitude (20 to 200) were compared alone, compressing the link confounder substantially. The LCP gradient shrank from 6.3 positions to 3.8, INP from 5.2 to 4.1, and CLS from 1.4 to 0.8. The gradients did not disappear, which argues against the pure-confounder story, and INP retained the largest share of its effect, consistent with interaction responsiveness carrying independent weight.

Second, sites were paired within verticals by keyword difficulty bands, so hard-keyword Good sites were not compared against easy-keyword Poor sites. The threshold shape survived: the Poor bucket still underperformed by 3 to 5 positions on average, while the Good versus Needs improvement gap stayed small.

The honest conclusion from the controls

After controlling for the measurable confounders, a residual association remains, concentrated in the Poor bucket and strongest for INP, then LCP, then CLS. Three interpretations fit this residual:

  1. Direct causation: Google’s page experience signals demote pages with Poor field vitals, exactly as the documentation describes.
  2. Indirect causation through behavior: Poor vitals degrade engagement (bounces, short sessions, abandoned interactions), and Google’s systems respond to the behavioral fallout rather than the metrics themselves.
  3. Unmeasured confounders: something else that correlates with both vitals and rankings, such as overall site maintenance quality, crawl efficiency, or technical SEO hygiene, is doing the real work.

The data cannot distinguish between these three, and any article claiming otherwise is overselling. But all three point to the same action: fix Poor vitals. If the effect is direct, the fix helps rankings. If behavioral, it helps engagement, which helps rankings and conversions regardless. If the real driver is general technical hygiene, performance work is part of that hygiene. The causation debate does not change the to-do list. The only scenario that changes it is chasing perfection within Good, where neither the data nor Google’s documentation shows any payoff.

The business case anyway: the conversion math, step by step

Rankings are only half the argument, and the weaker half. The stronger case for Core Web Vitals investment is the revenue that leaks through slow, janky pages every day, whether or not Google notices. Public benchmarks make the scale concrete. Google’s widely cited analysis of mobile page speed found that as load time rises from 1 second to 3 seconds, the probability of a bounce increases by 32 percent, and from 1 to 5 seconds it increases by 90 percent. Industry case studies point the same way: Walmart reported a 2 percent conversion increase for every 1 second of load-time improvement, and Mobify reported a 1.6 percent conversion lift per 100 ms. These are publicly reported benchmarks from the companies involved, cited here as directional evidence, not as guarantees for any specific site.

To show how this compounds, here is a worked example for a lead-generation site of the kind our lead generation services typically engage with. Every number below is illustrative, built from the public benchmarks above, so the arithmetic is transparent and checkable.

Monthly revenue = Sessions × Conversion rate × Close rate × Deal value 50,000 × 2.0% × 15% × $4,000 = $600,000 per month baseline

Step 1: establish the baseline. The site receives 50,000 sessions per month. At a 2.0 percent visitor-to-lead conversion rate, that yields 1,000 leads. With a 15 percent lead-to-customer close rate and a $4,000 average deal value, monthly revenue attributable to the site is $600,000.

Step 2: apply the speed improvement. LCP improves from 3.8 seconds (Poor) to 2.4 seconds (Good), a 1.4-second gain from image optimization and CDN caching. At the conservative end of the public benchmarks, each second of improvement lifts conversion by 1.5 percent in relative terms, so the rate rises approximately 2.1 percent relative: from 2.00 percent to 2.042 percent.

Step 3: compound through the funnel. At 2.042 percent, the same 50,000 sessions produce 1,021 leads instead of 1,000. At the unchanged 15 percent close rate, that is 153.2 customers instead of 150, and at $4,000 per deal, monthly revenue rises to $612,600. The gain is $12,600 per month, or $151,200 per year, from a one-time engineering effort.

Step 4: add the bounce-rate effect. The calculation holds sessions constant, which understates the gain: cutting load time from near 4 seconds to near 2.5 seconds reduces early exits, so engaged sessions rise too. If bounce-driven session loss falls by even 5 percent, the annual gain roughly doubles. Treat that as upside rather than a headline number, since it is harder to measure precisely.

The exact dollar figure is not the point; the structure is. Performance improvements compound through every downstream funnel stage, so modest relative lifts produce material revenue at scale. Rankings may or may not move when LCP drops by a second; revenue moves with far more certainty, because fewer frustrated users abandoning the page does not depend on any algorithm’s opinion.

What to fix first: the priority matrix

Given limited engineering time, the rational order is not “fix every metric” but “fix the failures that cost the most for the least effort.” The matrix ranks the most common fixes observed across the 100 sites by effort and expected impact, combining this study’s ranking data with publicly documented conversion effects.

FixTargetsEffortImpactNotes
Compress and resize above-the-fold images; serve AVIF/WebP with explicit width and heightLCP, CLSLowHighThe single most common root cause of Poor LCP in the sample. Explicit dimensions also prevent the layout shifts that inflate CLS.
Enable CDN caching and reduce server response time under 600 msLCPLow to mediumHigh27 of 31 Poor-LCP sites had slow TTFB. No front-end optimization compensates for a slow origin.
Defer non-critical JavaScript; split bundles; delay chat widgets and tag managersINPMediumHighINP showed the most consistent ranking gradient. Long main-thread tasks are the usual culprit; measure with the INP attribution build in Chrome DevTools.
Preload the LCP image and critical fonts; use font-display: swapLCP, CLSLowMediumCheap wins on the margin. Font swapping without reserved space causes shifts, so pair with size-adjust or fallback metrics.
Reserve space for ads, embeds, and dynamic content slotsCLSLowMediumFixes the worst CLS offenders. Ranking impact is small in isolation but the conversion impact on mobile commerce is real.
Remove or consolidate third-party scripts (heatmaps, redundant analytics, social widgets)INP, LCPMediumMedium to highThird parties were implicated in most Poor INP readings. Audit with request blocking to quantify each script’s cost before negotiating with stakeholders.
Full front-end rebuild or framework migrationAllHighUncertainOnly justified when the stack itself is the bottleneck. The data gives no reason to rebuild a site that already scores Good.

Two rules of thumb: work the Poor bucket before touching anything else, and do not optimize within Good. Hours spent shaving a 1.9-second LCP to 1.4 seconds are hours not spent on content, links, or conversion work that demonstrably moves the needle.

FAQ

Does Google use lab scores from Lighthouse in rankings?
No. Google’s ranking systems use field data from the Chrome User Experience Report, aggregated over a 28-day rolling window. Lighthouse and WebPageTest produce lab data: useful for diagnosis, but the lab number never enters the ranking calculation. A page can score 100 in Lighthouse yet show Poor vitals in Search Console. Optimize against CrUX field data, not the lab score.
Is there any ranking benefit to perfect scores once a page is already Good?
According to Google’s documentation, no. Search Central states that reaching the Good thresholds is sufficient, with no additional ranking benefit from further optimization. This study agrees: within the Good bucket, exact metric values showed no relationship to rankings. Perfect scores are a fine engineering goal, not an SEO strategy.
Is Core Web Vitals a major ranking factor?
The honest characterization is that Core Web Vitals act as a quality floor, with most of the effect concentrated in the Poor bucket. Content relevance, link authority, and intent match remain the primary ranking drivers by a wide margin. Treat vitals as a tiebreaker among comparable pages, plus a penalty for broken experiences, not as a lever that lifts good pages higher.
Why did INP correlate more consistently than LCP?
INP captures page behavior across the whole visit, not just during load. Poor-INP pages were typically overloaded with JavaScript: heavy bundles, competing tag managers, third-party widgets blocking the main thread. Those pages tend to be worse in broader, harder-to-measure ways, so INP functions as a summary indicator of overall sloppiness. LCP, by contrast, can be fixed with one image optimization while the rest of the page stays poor, which weakens its signal.
Can good Core Web Vitals compensate for weak content?
No, and Google says so explicitly: good page experience does not override great, relevant content. Several all-Good sites in this sample ranked poorly on thin or mismatched content, while some Needs improvement sites ranked well on exceptional content and links. Performance multiplies a foundation built by content and authority; it does not substitute for either.
How long after fixing vitals should ranking changes appear?
Plan for 4 to 8 weeks minimum. CrUX is a 28-day rolling aggregate, so a fix only fully enters the dataset after a month of real visits, and low-traffic pages take longer to accumulate URL-level data. In this study, sites that improved INP showed average gains about four weeks after deployment. Treat this work as infrastructure, not a campaign.

Conclusion

So, do Core Web Vitals affect rankings? The careful answer from 100 sites and 60 days of data: yes, but as a floor, not a ladder. Pages with Poor field vitals ranked measurably worse, with the effect strongest for INP, visible for LCP, and weak for CLS. Once pages reached Good thresholds, the ranking relationship disappeared. Causation remains genuinely open, tangled with content, links, and brand, but every plausible interpretation points to the same action: fix Poor vitals, stop at Good, and spend the rest on content and links.

The more important finding has nothing to do with Google. Slow, unresponsive pages leak revenue through bounces and abandoned interactions at a scale that dwarfs most ranking effects, governed by user psychology rather than any algorithm. The conversion math here is illustrative, but its direction is among the best-documented facts in web performance research. Performance work pays for itself even if rankings never move.

For teams that want this handled properly, the sequence is straightforward: audit field data in Search Console and PageSpeed Insights, fix the Poor bucket using the priority matrix above, verify against CrUX rather than lab scores, then move on. If that sounds like work better delegated to specialists, SCORSH’s SEO services include technical performance audits as part of every engagement, and the team works with businesses across India and the US on exactly this kind of compounding, unglamorous growth work.

Rishabh
Rishabh

Rishabh is the founder of SCORSH, a performance marketing agency working with businesses across India and the US. With 14 years of experience, he writes about SEO, paid media, and the math behind growth.