Understanding load balancing: what it does and how it decides
Quick answer
A load balancer sits in front of a group of servers and spreads incoming traffic across them, so no single server carries the whole load and the service stays up if one of them fails. It decides where each request goes using health checks to confirm a backend is actually responding, then an algorithm such as round robin or least connections to pick which healthy backend gets the request.
What a load balancer actually does
Run a service on one server and that server is both your capacity ceiling and your single point of failure. A load balancer solves both problems at once. It sits in front of two or more backend servers, all running the same application, and distributes incoming requests across them. Traffic that would have overwhelmed one server gets spread across several, and if one backend goes down, the load balancer stops sending it traffic and the others keep serving requests without the visitor noticing.
The backends themselves don't need to know a load balancer is there. As far as they're concerned, requests just arrive. The coordination, deciding which backend handles which request, happens entirely at the load balancer.
Health checks: only sending traffic to backends that are actually up
A load balancer that blindly cycles through a fixed list of servers is only half useful, because a server that's crashed or hung still sits in that list. Health checks close that gap. The load balancer polls each backend on a regular interval, typically a simple TCP connection attempt or an HTTP request to a defined path, and marks a backend as unhealthy the moment it stops responding correctly. Unhealthy backends are removed from rotation until they start passing checks again.
This is what makes a load balancer resilient rather than just a traffic splitter. Without health checks, a single failed backend would keep receiving a share of requests and every visitor routed to it would see an error. With them, failure in one backend is invisible to visitors, as long as the remaining backends have enough spare capacity to absorb the difference.
Common algorithms: how the next request gets assigned
Once the load balancer knows which backends are healthy, it still needs a rule for picking which one gets the next request. Two approaches cover most real deployments:
- Round robin cycles through the healthy backends in turn: first request to server A, next to server B, next to server C, then back to A. It's simple and works well when requests are roughly similar in cost and backends are similar in capacity.
- Least connections sends each new request to whichever healthy backend currently has the fewest active connections. This adapts better than round robin when requests vary widely in how long they take to handle, since it avoids piling more work onto a backend that's already busy with slow requests.
Other algorithms exist, weighted variants that favour more powerful backends, or ones that route based on source IP, but round robin and least connections cover the large majority of setups.
Sticky sessions: keeping a visitor on the same backend
Spreading requests across backends works cleanly when each request is independent. It breaks down when an application keeps session state, a shopping cart, a login session, in memory on whichever server first handled that visitor. If the next request from the same visitor lands on a different backend, that state isn't there and the visitor gets logged out or loses their cart.
Sticky sessions solve this by pinning a visitor's requests to the same backend for the life of their session, usually by setting a cookie the load balancer reads on each subsequent request. It's a workaround rather than a fix: the underlying issue is session state living on one server instead of somewhere shared, such as an external cache or database, which is the more robust way to design around this once an application needs to scale beyond one backend. Sticky sessions also slightly undermine even load distribution, since a backend that happened to pick up several long-running sessions stays busier than one that didn't.
TLS termination: where encryption gets handled
When traffic to the service is encrypted, the load balancer has to decide what to do with that encryption, and there are two common patterns:
- TLS passthrough: the load balancer forwards encrypted traffic to the backend untouched, and the backend itself terminates the TLS connection. The load balancer never sees the decrypted content, and each backend needs its own certificate configuration.
- TLS offloading: the load balancer terminates the TLS connection itself, decrypting incoming traffic, then forwards it to the backend either unencrypted or re-encrypted on a separate connection. This centralises certificate management in one place and takes the computational cost of encryption off the backends, at the cost of the load balancer seeing decrypted traffic, which matters for anything sensitive to inspect.
Which pattern fits depends on where you need certificate management to live and whether backends can afford the overhead of terminating TLS themselves. Neither is universally correct.