What really moves PageSpeed Insights: caching, hosting and your stack

15 min read

The score is the visible end of architecture decisions. Caching, rendering, stacks and hosting compared, and the business results speed has bought.

PageSpeed Insights explained

In 2021, Vodafone Italy ran an A/B test on one of its landing pages. The two versions looked identical and did the same things. In one of them, the team had moved part of the rendering from the browser to the server and stopped loading images nobody could see. On real visitors' phones, the main content of that version appeared after 5.7 seconds instead of 8.3. That version made 8% more sales.

Vodafone didn't get there by tweaking the page until a score turned green. Its faster version still didn't meet Google's own "good" threshold. What changed was an architectural decision: where the page was rendered and what the browser had to wait for.

That's what this article is about. PageSpeed Insights hands you one number from 0 to 100, and it's tempting to treat it as a grade to be improved line by line. In practice, most of that number is decided long before anyone opens the report: by what is cached and where, by where the page is rendered, by the stack the site is built on and by the platform it runs on. Those are the decisions worth understanding, and the ones that show up in revenue.

Reading the report in one minute

PageSpeed Insights shows two different things.

At the top is what real visitors experienced: field data from Chrome users over the last 28 days, at the 75th percentile. That means it describes your slower visitors, on older phones and weaker connections, and that a fix deployed today takes weeks to show fully. Google's ranking systems use this part.

Below it is one simulated test, run by Lighthouse on an emulated mid-range phone. The 0 to 100 score belongs to this part. It's useful for finding causes, but it isn't what Google ranks and it isn't what your visitors see.

The field data covers three metrics, the Core Web Vitals. Each one points at a different layer of the system:

  • LCP (Largest Contentful Paint) is when the main content appears. Good is 2.5 seconds or less. It's mostly decided by the server, caching and delivery: how fast the first byte arrives and how soon the browser can fetch the main image. Tokopedia, an Indonesian marketplace, improved it by 55% and saw sessions grow 23% longer. Nykaa, an Indian retailer, improved it by 40% and got 28% more organic traffic.
  • INP (Interaction to Next Paint) is how quickly the page responds to a tap or a click. Good is 200 milliseconds or less. It's decided by JavaScript: how much of it runs in the browser, and when. redBus, a bus-ticketing site, reworked a few heavy interactions, such as a scroll handler and a form that re-rendered half the page on every keystroke, and saw sales rise by 7%.
  • CLS (Cumulative Layout Shift) is how much the content jumps around. Good is 0.1 or less. It's decided by layout discipline: reserving space for images, ads and banners before they arrive. iCook, a recipe site, reserved fixed slots for its ads. Fewer slots were filled, and ad revenue still rose by 10%, because ads that stay in place get seen.

And the effect isn't limited to large changes. A Deloitte study for Google of mobile sites from 37 brands found that a 0.1-second improvement went with 8.4% more conversions on retail sites.

Caching: where most of the speed actually comes from

The fastest request is the one that never reaches your server. Every layer of caching between the visitor and the database removes work, and most slow sites are missing at least one of them.

  • Browser cache. Repeat visits and the next page download nothing they already have. Typical gap: files served without cache headers, so the browser asks again every time.
  • CDN (edge) cache. Files and pages are served from a location near the visitor, without touching your server. Typical gap: only images and scripts are cached at the edge, while every page is still generated on request.
  • Full-page cache. A finished page is stored and served as-is to everyone who sees the same thing. Typical gap: pages that set a cookie on every response can't be cached at all.
  • Application and data cache (Redis, Memcached). Expensive queries and calculations are done once, not on every request. Typical gap: the same query repeated dozens of times per page, or caching the wrong things.
  • Runtime cache (OPcache in PHP, a warm process elsewhere). Code isn't recompiled or reloaded for every request. Typical gap: a cache that is too small, or cold starts after every quiet period.

Two of these gaps are worth explaining, because they're common and invisible.

A cookie can make a page uncacheable. CDNs and page caches won't share a response that sets a cookie, because the cookie might belong to one visitor. Many frameworks start a session on every page by default, even for anonymous visitors who will never log in. Laravel's standard web routes set a session cookie and a CSRF cookie on every response. PHP's own sessions send headers that forbid caching. The result is a blog or a product page, identical for every visitor, that is generated from scratch on every single request.

Cache lifetime is a decision about how files change, not about time. The obvious instinct is to keep caching short, so that changes appear quickly. The better approach is the opposite: cache files for a year, and change their address when their content changes. Modern build tools already do this for scripts and stylesheets, by putting a fingerprint of the content in the file name. Uploaded images can do the same with a version in the URL. Then nothing is ever stale, and nothing is ever downloaded twice. Short lifetimes look safer, but they make every visitor ask the server again and again, and the files still aren't fresh if a CDN in between keeps the old copy.

I found both kinds of gap on this site while writing this article. Uploaded images were served without any cache headers, so every image transformation started a server function that read the original file from storage again. And a CDN proxy in front of the hosting platform rewrote cache lifetimes to four hours, on top of the platform's own CDN. Neither problem showed up in the score. Both showed up in the headers.

Where the page is rendered

The second big decision is where the HTML is produced. Each option trades something:

  • At build time (static generation). Pages are generated once and served from the CDN. The fastest option, and the cheapest to run. It only works for content that is the same for everyone, and content changes need a rebuild or a scheduled refresh.
  • On the server, cached (incremental regeneration or a full-page cache). Pages are generated on the first request and then served from cache until the content changes. Nearly as fast as static, and it works for sites with many pages and frequent updates. It needs a reliable way to refresh the cache when content changes, and that is where most bugs live.
  • On the server, on every request. Always fresh, and needed for personalised pages, but every visitor pays the full time of the server and the database.
  • In the browser (single-page applications). The server sends an almost empty page, and JavaScript builds the content after it has downloaded and run, often after fetching the data as well. The main content appears last by design. This was Vodafone's problem, and it is the most common reason a modern site fails LCP.

Most sites need a mix: static or cached pages for content and marketing, server rendering for accounts and checkout, and as little browser rendering as the interface allows.

The stack: what you get for free and what you have to watch

No stack is fast or slow on its own. Each one makes some things easy and some mistakes likely.

  • WordPress. Fast by default: mature page caching, with plugins or at the host. Usually goes wrong: heavy themes and page builders, and every plugin adding its own scripts to every page. The plugins that slow a site down are rarely the ones that were meant to speed it up.
  • Shopify. Fast by default: hosting, CDN and image delivery are handled, and handled well. Usually goes wrong: apps inject scripts into every page, and uninstalling an app doesn't always remove its code. There is little control over the server.
  • Laravel and other PHP frameworks. Fast by default: server-rendered HTML, so the main content arrives with the first response. Usually goes wrong: sessions and cookies on every page prevent full-page caching, and server time depends on OPcache, queues and database queries.
  • Next.js and similar frameworks. Fast by default: static and cached pages, image optimisation and code splitting are built in. Usually goes wrong: it's easy to make a page dynamic by accident, for example by reading cookies, and then nothing is cached. Large client-side components still cost INP.
  • Single-page applications (React, Vue or Angular rendered in the browser). Fast by default: quick, smooth navigation once loaded. Usually goes wrong: the first view waits for the JavaScript and the data, and server rendering usually has to be added later.
  • A headless CMS with a front-end framework. Fast by default: content and presentation are separate, so pages can be static or cached. Usually goes wrong: publishing has to refresh the right caches. If it doesn't, editors publish and nothing changes.

The last point is a lesson I learned on this site as well. Publishing an article from the admin of a local copy didn't refresh the cache of the live site, because the refresh only reaches the host it was triggered from. The page was fast, and it was also out of date.

The platform: where the site runs matters as much as what it's built with

The same application behaves very differently depending on where it's deployed.

Shared hosting. Cheap and simple. Many sites share one server, so performance depends on your neighbours, and you rarely control caching or server settings. Fine for small sites, but server response time is often the limit.

A VPS or dedicated server. Full control. You can set up a page cache in front of the application, Redis, a tuned PHP worker pool and exactly the cache headers you want. The price is that you own all of it: updates, security, monitoring, and a CDN if you want one.

A managed platform for one framework (such as Laravel Cloud, or a PaaS). Good defaults and less operational work, with scaling handled for you. Less control over the details, and the cost rises with traffic.

Serverless and edge platforms (such as Vercel or Netlify). A global CDN, static and cached pages and image optimisation come built in, and for a framework like Next.js it's often the fastest setup with the least effort. Three things need attention:

  • Region. Server functions run in one region by default, and the database may be somewhere else. This site's functions ran in the eastern United States while its database was in Ireland, so every query crossed the Atlantic. Moving the functions next to the database was one line of configuration.
  • Cold starts. A function that hasn't run for a while takes longer to answer, which on a low-traffic site can be most visits.
  • Usage-based pricing. Image transformations, function time and bandwidth are billed per use, so caching decisions are also cost decisions.

Hosted commerce and site builders (such as Shopify). Infrastructure is solved for you. Performance depends almost entirely on the theme, the apps and the third-party scripts you add.

A CDN or proxy in front. On a VPS, a CDN such as Cloudflare is one of the best investments available. In front of a platform that already has its own CDN, it adds a second cache that can rewrite headers and serve stale copies, which makes the system harder to reason about. Choose one layer to own caching, and let the other pass through.

The gaps that cost the most

Most slow sites aren't slow because of one bad line. They are slow because of a few structural gaps, and the same ones appear in almost every stack:

  1. Latency between tiers. Every request that crosses a network pays a round trip: between the visitor and the server, the application and the database, the application and its cache, the application and external APIs. Inside one data centre a round trip takes well under a millisecond. Across an ocean it takes tens of milliseconds. A page that runs its queries one after another, or fetches related records one by one (the N+1 pattern), pays that cost dozens of times before it sends its first byte. Distance multiplied by the number of sequential calls is often the largest part of server response time, and no amount of code optimisation removes it.
  2. Responses that can't be cached. A page that is the same for every visitor is still generated on every request if its response is marked as private: a Set-Cookie header, Cache-Control: private or no-store, or a Vary on cookies. Caches also fragment on their key. Marketing parameters such as utm_source create a separate cache entry for every campaign link unless the CDN is told to ignore them, so the visitors who arrive from paid campaigns are often the ones who get uncached pages.
  3. Missing or wrong freshness rules. Without Cache-Control, browsers and CDNs guess or revalidate on every request. Short lifetimes turn each repeat view into a conditional request, which saves the download but still costs a round trip per file. Long lifetimes on addresses that don't change with the content serve stale files that can't be recalled. The robust combination is a long lifetime on versioned URLs, so that freshness comes from the address, not from the timer.
  4. A long critical rendering path. The main content can only appear after everything it depends on has arrived. Rendering in the browser puts it behind a JavaScript download, its execution and usually an API call. Render-blocking stylesheets and scripts delay the first paint, and a main image referenced only from CSS or inserted by script is discovered late. Each dependency in the chain adds at least one round trip to LCP.
  5. Contention on the main thread. The browser runs JavaScript, handles input and renders the page on one main thread. Framework hydration, large bundles and third-party scripts (tag managers, chat widgets, A/B testing, consent platforms) compete for it, and any task longer than 50 milliseconds delays the response to whatever the visitor just tapped. Third-party code is the hardest part to control, because it's added through admin panels and tag managers rather than code review, and it grows over time.
  6. Asynchronous content without reserved space. Images without dimensions, ads, embeds, banners injected after load and web fonts that change the size of text all move content that has already been painted. Layout stability is decided when the page is designed, by reserving space for everything that arrives later.
  7. Cold paths. A serverless function that hasn't run recently starts slowly. A CDN cache is empty for every new region, page and image variant. A deploy can reset runtime caches. On a busy site these costs are spread thin. On a low-traffic site, a large share of real visits hits a cold path, while the developer's own repeat visits and the lab test, run against warm caches, never see it.
  8. Measuring the wrong population. A lab test on a simulated phone, or a check on a fast laptop on office fibre, describes one load under one set of conditions. Field data at the 75th percentile describes the slower quarter of real visits, on older devices and weaker networks. Decisions made from the first kind of data routinely miss the problems that cost the most in the second.

None of these is fixed by tweaking a tag. Each is an architectural or operational decision, and each one shows up in response headers, server timings or real-user data long before it shows up in the score.

Measure what visitors experience

The companies in this article all did the same thing: they measured real visits and fixed specific causes.

  • Google Search Console groups your pages by Core Web Vitals problem, using the same field data as PageSpeed Insights.
  • Real-user monitoring measures every visit and tells you which element or interaction was slow. Hosting platforms such as Vercel include it, and Google's web-vitals library, which redBus used, works anywhere.
  • Response headers show what the score can't: whether a page or a file came from a cache, which cache, and for how long it may be kept. Reading them is often the fastest way to find the gaps above.

The score is the visible end

Vodafone rendered differently and sold 8% more. redBus changed how a few interactions worked. iCook reserved space for its ads. On this site, the gaps were a missing header, a proxy rewriting cache lifetimes and functions in the wrong region. None of those decisions appears in a number between 0 and 100, but all of them move it.

Use the score to find where to look. Then look at the architecture behind it: what is cached and where, where pages are rendered, what the stack does by default and what the platform does on top. That's where the speed is decided, and where the business results come from.


Sources: Vodafone, redBus and The business impact of Core Web Vitals (Tokopedia, Nykaa, iCook) on web.dev; Milliseconds Make Millions, Deloitte for Google.