What Is a Load Balancer?
Picture a bank with one teller. Every customer who walks in joins the same line, and the teller serves them one at a time no matter how busy it gets. Now picture that bank opening two more teller windows — but customers still don't know which one is free, so everyone keeps joining teller 1's line while windows 2 and 3 sit empty. Adding tellers didn't actually help, because nothing is directing customers toward the free ones.
Now add a queue manager standing at the door. A new customer walks in, and the queue manager looks down the row of tellers and sends that customer to whichever one can help them fastest. No customer has to guess which line is short. No teller sits idle while another one drowns. That queue manager is a load balancer — it never serves a customer itself, it only decides who goes where.
A load balancer doesn't create capacity — it reveals it. Three idle tellers already existed in both banks above; only the second bank could actually put them to use.
Swap “customer” for a request, “teller” for a backend server, and “queue manager” for a piece of software (or a dedicated appliance) sitting in front of your fleet, and that's the whole idea. Everything else in this guide is really answering two questions the queue manager has to solve on every single request: which backend is actually available right now, and which one should this particular request go to.
What a Load Balancer Does
A load balancer is a network service placed in front of backend servers. Clients connect to the load balancer rather than choosing an application server directly. The load balancer then selects an eligible backend and forwards the work.
This indirection creates one stable entry point while allowing the backend fleet to grow, shrink, or lose unhealthy members. It does not create backend capacity by itself. It makes existing capacity reachable and controllable.
A service that accepts incoming connections or requests and forwards each one to an eligible backend.
The load balancer endpoint that accepts traffic on a protocol and port, such as HTTPS on port 443.
A server instance that can receive forwarded work from the load balancer.
A named or logical group of backends eligible for a particular class of traffic.
A condition and action that decide which backend pool receives traffic.
A periodic probe used to decide whether a backend should remain eligible for new traffic.
Continued service after a component fails, using healthy capacity that remains available.
A connection currently open on a backend. Duration determines how long it contributes to current work.
A relative value that controls how much traffic one backend receives compared with others.
A common proxy term for the backend service or pool to which traffic is forwarded.
The forwarding path
A valid path has three stages: traffic arrives, the load balancer makes a routing decision, and a healthy backend handles the work. A server drawn on the canvas but not connected to the load balancer contributes no usable capacity.
“A load balancer is what gives a system more capacity.”
It doesn't. A load balancer distributes traffic among instances that already exist. An autoscaler changes the number of instances. Many systems use both, but they solve different problems — a load balancer with three idle backends behind it has exactly as much capacity as those three backends provide, no more.
Choose What the Load Balancer Needs to Understand
Layer 4 and Layer 7 refer to layers in the network model. The distinction is entirely mechanical: a Layer 4 load balancer forwards a transport connection using addresses and ports and never opens it, so it cannot see anything about the HTTP request inside. A Layer 7 load balancer terminates the connection and reads the actual HTTP request — method, path, headers, cookies — before it decides where to send it.
Forwards the TCP/UDP connection. Never opens the request itself.
Terminates the connection and reads the actual HTTP request.
| Question | Layer 4 | Layer 7 |
|---|---|---|
| What can it inspect? | IP, port, TCP or UDP connection | HTTP host, path, method, headers |
Can it route /api and /media differently? | No | Yes |
| Typical use | General TCP services, very high throughput | HTTP APIs, host and path routing, request policies |
TLS Termination & the Reverse Proxy Role
A load balancer sitting in front of HTTPS traffic almost always plays a second role: it terminates TLS. The client's encrypted connection ends at the load balancer, not at the backend. The load balancer holds the TLS certificate and private key, decrypts the incoming request, and only then can a Layer 7 routing rule actually read the path, headers, or cookies described in the previous chapter — encrypted bytes can't be inspected for a path match.
Traffic between the load balancer and the backend is commonly plain HTTP again, since it stays inside a private network. Some deployments re-encrypt that hop for defense in depth, but either way the certificate lives in exactly one place — the load balancer — instead of being copied onto every backend instance.
Why keep the certificate in one place?
Choose the Signal That Represents Backend Work
A routing algorithm is the rule used to select a backend. The correct algorithm depends on whether servers are equal, whether request cost varies, and whether current connection count is a useful signal.
Equal request count. Best for similar servers and similar request cost.
Configured traffic ratio. Best when backend capacities differ predictably.
Current open connection count. Best when connection durations differ.
Cycles through eligible backends in turn. It assumes each request costs roughly the same and each backend has roughly the same capacity.
Cycles according to configured weights. Deterministic, and appropriate when backend capacities differ in known proportions.
Chooses the backend with the fewest active connections. It reacts to long-lived requests that remain open while short requests finish.
“A load balancer just splits traffic evenly, no matter what.”
Round Robin does — but it's a deliberate default for a uniform fleet, not the definition of load balancing itself. Weighted Round Robin and Least Connections exist specifically because backends are often not identical and requests are often not equal-cost. Picking “even” when the fleet is heterogeneous overloads the smaller or slower-clearing backends while the rest sit under-used.
A fleet has one Large backend and three Small backends, where Large has 4x the rated capacity of a Small. Which algorithm keeps every backend at roughly the same percentage of its own capacity?
Route in Proportion to Backend Capacity
Round Robin sends equal request counts. That is unsafe for a mixed fleet because a Small server and a Large server receive the same traffic despite having different limits. Weighted Round Robin lets the routing share mirror capacity.
“Load-balancer weights have to add up to 100, like percentages.”
Weights 22, 43, 80, and 80 sum to 225 in the figure above. The load balancer normalizes them into traffic shares automatically — they only need to represent relative capacity, not a literal percentage breakdown.
Stop Routing Before Spare Capacity Can Help
A failed backend must be removed from the eligible pool. Health checks perform this control-loop function by probing each backend and changing its routing state when failures cross a threshold.
“Health checks alone guarantee zero failed requests during a backend failure.”
There is always a detection window — the time between a backend actually failing and the health check crossing its failure threshold and removing it from rotation. Any request routed to that backend during the window still fails. Faster checks shrink the window; they cannot eliminate it.
Health detection and failover capacity are separate
Health checks prevent requests from reaching a failed server. They do not create replacement capacity. After removal, the remaining servers must still handle the specified failure traffic without crossing their utilization limit.
What makes a useful health check
Probe a path that proves the service can perform meaningful work, not only that a process exists.
The probe runs frequently, so it must not create significant application or database load.
Detection must be quick enough for the recovery objective without reacting to one harmless transient error.
Avoid a probe that declares every backend unhealthy because one optional shared dependency is degraded.
Graceful Deregistration & Connection Draining
Health checks answer whether a backend is eligible for new traffic. They don't answer what happens to the requests already in flight on that backend the moment it's marked unhealthy, or simply scheduled for planned removal — a deploy, a scale-down, a maintenance window.
A hard removal stops sending requests instantly and drops any connection already in progress mid-response. Connection draining (also called graceful deregistration) stops sending new requests immediately, but lets requests already in flight finish normally, up to a configured drain timeout. Only once every in-flight request completes — or the timeout expires — is the backend fully removed.
The backend is removed instantly. Anything already running on it is cut off.
New traffic stops right away, but existing requests get to complete before removal finalizes.
Why does this matter for deploys, not just failures?
Request Count Is Not Current Work
Round Robin equalizes arrivals, not active work. When all requests finish in similar time, that distinction may not matter. When some requests remain open far longer, one backend can accumulate more active work than another even after receiving the same number of requests.
Many requests complete quickly, so each connection disappears almost immediately.
Only a few exports arrive each second, but they remain open long enough to dominate active connection count.
Round Robin distributes arrivals evenly. It cannot distribute active work evenly, because arrivals and active work are only the same thing when every request takes the same amount of time.
Least Connections uses open connection count as a live approximation of backend work. It does not replace capacity planning. The fleet still needs enough RPS capacity, and connection count is only useful when it correlates with resource usage.
Route Each Workload to Dedicated Capacity
A Layer 7 routing rule can inspect the URL path and choose a backend pool. This allows one domain to serve multiple workloads while keeping their compute capacity independent.
Pool isolation
A backend is dedicated only when it belongs to one pool. Connecting the same application server to two route pools creates shared capacity and allows one route to affect the other. Isolation requires separate membership, separate sizing, and separate utilization checks.
Why does the practice canvas show a separate load-balancer node per path?
A Repeatable Interview Method
Separate topology, algorithm, health, and capacity decisions. A correct answer should explain what traffic is being routed, which signal drives selection, which backends are eligible, and what happens when one fails.
- 1
Identify the traffic unit
State whether the input is requests, connections, jobs, bytes, or sessions. Keep units attached to every calculation.
- 2
Choose Layer 4 or Layer 7
Use Layer 7 when routing depends on HTTP host, path, method, or headers. Use Layer 4 when transport-level forwarding is sufficient.
- 3
Decide where TLS terminates
If the routing rule needs to read the request, TLS has to terminate before that decision is made.
- 4
Define backend pools
List which servers may handle each traffic class. Decide whether capacity is shared or isolated.
- 5
Choose the routing algorithm
Use Round Robin for a uniform fleet, Weighted Round Robin for known capacity differences, and Least Connections for varying connection duration.
- 6
Calculate normal capacity
Verify aggregate throughput and every backend's utilization under the selected traffic split.
- 7
Define health and eligibility
Specify how failed backends are detected and removed from rotation, and how in-flight requests are drained on planned removals.
- 8
Calculate failure capacity
Remove the failed backend from the math and verify that the remaining fleet serves the required failover load.
- 9
Check isolation and cost
Verify that no backend is unintentionally shared and that every load balancer and server fits the monthly budget.
Do not trust aggregate capacity
A bad traffic split can overload one backend while total fleet capacity looks sufficient.
Do not ignore duration
Equal arrival counts do not imply equal active work when requests have different lifetimes.
Do not fake isolation
A backend shared by two pools is still a shared failure and capacity domain.
Apply Each Traffic-Routing Concept
Each lab isolates one routing decision. Build mode asks you to construct the system from a specification. Fix mode gives you a concrete routing failure and asks for the smallest effective repair.
Uneven Fleet
Use Weighted Round Robin to distribute traffic across backends with different capacities.
The Dead Server Still Gets Traffic
Enable health checks and prove the remaining fleet can carry failover traffic.
Long Requests Break Round Robin
Use Least Connections when active connection duration matters more than arrival count.
You are ready when you can justify the routing layer, algorithm, health behavior, backend membership, failure capacity, and cost without relying on the diagram alone.