HTTP caching with Varnish: speeding up a web server without changing it
Quick answer
Varnish is a reverse proxy cache. It sits in front of your web server and stores copies of responses, so a repeat request for the same content is served straight from memory instead of reaching the application behind it. That's faster for the visitor and lighter load on the origin server, and it requires no changes to the application itself.
What Varnish actually does
Varnish sits between the internet and your web server. A request for a page arrives at Varnish first. If Varnish already has a valid cached copy of that exact response, it returns it immediately, the request never reaches your application, your database, or whatever else sits behind it. If Varnish doesn't have a cached copy, or the copy it has has expired, it fetches a fresh one from the origin server, serves it to the visitor, and stores a copy for next time.
This is a reverse proxy, not a browser cache. A browser cache only helps the same visitor on a repeat visit. Varnish helps every visitor, because the cached copy lives on the server side and is shared across all of them. It's also not a CDN in the sense of distributing content across many geographic locations, Varnish typically runs as a single layer close to your origin server. See What is a CDN, and how content delivery networks work for how that compares to caching spread across edge locations worldwide.
What it's naturally good at
Varnish works best on content that looks the same for every visitor: a static marketing page, a blog post, a product listing on an e-commerce site, a JSON response from an API endpoint that doesn't vary by who's asking. None of that content depends on who's making the request, so one cached copy can safely serve thousands of different visitors. This is exactly the kind of workload that benefits most: pages that are expensive to generate (a database-backed product catalogue, for example) but don't change on every request.
The gain compounds under load. A page that takes, say, a few hundred milliseconds to render from the application only has to be rendered once per cache period, not once per visitor. Every other request in that window is served from memory in a fraction of the time, and the origin server barely notices the traffic.
The real complication: content that isn't the same for everyone
The difficulty with any caching layer shows up the moment content stops being identical for every visitor. A logged-in user's account page, a shopping cart, a page that shows "Welcome back, [name]", none of that can be cached and served to the next visitor without leaking one person's content to another. Get this wrong and the consequences aren't just a stale page, they're a serious privacy problem: one visitor seeing another visitor's session content.
Varnish doesn't guess at this on its own. It needs to be told, through its configuration language (VCL, Varnish Configuration Language), which requests are safe to cache and which aren't. Common patterns include:
- Never caching requests that carry a session cookie or an authentication header.
- Never caching specific paths, such as
/account/,/checkout/, or/wp-admin/. - Stripping cookies from the request before checking the cache for content that genuinely is the same for everyone, so an unrelated tracking cookie doesn't accidentally prevent caching that would otherwise be safe.
- Varying the cached response by a specific header, such as
Accept-Language, when the same URL legitimately returns different content depending on that header.
A basic VCL rule to bypass the cache for anything carrying a session cookie looks roughly like this:
sub vcl_recv {
if (req.http.Cookie ~ "session_id") {
return (pass);
}
}
return (pass) tells Varnish to fetch straight from the origin and not store the response, exactly what you want for anything personalised. Getting this rule set right for a given application, especially one that wasn't designed with a cache in front of it, is usually the bulk of the setup work.
The classic hard problem: cache invalidation
Even for content that's genuinely safe to cache, there's a second problem: making sure Varnish drops or refreshes a cached copy when the underlying content actually changes. If a product's price is updated in the application but Varnish is still serving yesterday's cached page, visitors see stale data until that cache entry expires or is explicitly cleared.
There are two broad approaches, and most real setups combine both:
| Approach | How it works | Trade-off |
|---|---|---|
| Time-based expiry (TTL) | Each cached response is kept for a set period, then automatically treated as stale and re-fetched from the origin on the next request | Simple to configure, but there's always a window where visitors can see outdated content |
| Explicit purging | The application (or an operator) tells Varnish directly to drop a specific cached entry the moment the underlying content changes, typically via an HTTP PURGE request to the affected URL | Content is fresh immediately, but it means wiring up a purge call from wherever content gets updated |
This is often described as one of the two genuinely hard problems in computing, and it's worth taking seriously rather than treating as an afterthought. A short TTL reduces the risk of stale content but also reduces how much load Varnish actually takes off the origin, since cached copies expire and get re-fetched more often. A longer TTL gives better cache efficiency but a longer window of potential staleness unless purging is wired up properly. There's no universally correct answer, it depends on how often the content in question actually changes and how much staleness is acceptable for that particular page.
Where Varnish fits with the rest of the stack
Varnish operates at the HTTP layer, in front of whatever serves your application, whether that's a plain web server, a WordPress or WooCommerce install, or a custom application server. See Hosting WordPress and WooCommerce on a dedicated server for a workload where this pattern comes up often, product and category pages are prime caching candidates, cart and checkout pages are not.
Because Varnish is doing extra work of its own, holding cached objects in memory and evaluating VCL rules on every request, the server it runs on still needs enough CPU and RAM to do that comfortably alongside the origin application, particularly under high request volume. See Choosing a CPU for CI/CD, virtualisation and Kubernetes nodes for general guidance on matching compute resources to a workload's actual demands.