Section 02 · Foundation

Load Balancers & Traffic Routing

Starting from a plain-English analogy for what a load balancer even is, build up to Layer 4 vs. Layer 7 routing, TLS termination, every major routing algorithm, health checks, graceful failover, and isolated backend pools.

By the end

You will be able to choose a routing layer and algorithm, calculate traffic shares, plan capacity after failure, and design isolated backend pools for different routes.

12 chapters·11 figures·4 labs
01
Before any of the terminology

What Is a Load Balancer?

Picture a bank with one teller. Every customer who walks in joins the same line, and the teller serves them one at a time no matter how busy it gets. Now picture that bank opening two more teller windows — but customers still don't know which one is free, so everyone keeps joining teller 1's line while windows 2 and 3 sit empty. Adding tellers didn't actually help, because nothing is directing customers toward the free ones.

Now add a queue manager standing at the door. A new customer walks in, and the queue manager looks down the row of tellers and sends that customer to whichever one can help them fastest. No customer has to guess which line is short. No teller sits idle while another one drowns. That queue manager is a load balancer — it never serves a customer itself, it only decides who goes where.

One teller, no queue manager
Customers
Everyone joins the same line, out of habit
Teller 1
Works through the line alone
Teller 2
Idle
Teller 3
Idle
A queue manager directs every customer
Customers
Arrive and wait to be directed
Queue manager
Sends each one to whichever teller is free
Teller 1
Busy
Teller 2
Free
Teller 3
Free
Figure 1. A load balancer is the queue manager, not another teller. It never does the underlying work itself — it only decides which teller does.
Key idea

A load balancer doesn't create capacity — it reveals it. Three idle tellers already existed in both banks above; only the second bank could actually put them to use.

Swap “customer” for a request, “teller” for a backend server, and “queue manager” for a piece of software (or a dedicated appliance) sitting in front of your fleet, and that's the whole idea. Everything else in this guide is really answering two questions the queue manager has to solve on every single request: which backend is actually available right now, and which one should this particular request go to.

02
Load balancer fundamentals

What a Load Balancer Does

A load balancer is a network service placed in front of backend servers. Clients connect to the load balancer rather than choosing an application server directly. The load balancer then selects an eligible backend and forwards the work.

This indirection creates one stable entry point while allowing the backend fleet to grow, shrink, or lose unhealthy members. It does not create backend capacity by itself. It makes existing capacity reachable and controllable.

Load balancer

A service that accepts incoming connections or requests and forwards each one to an eligible backend.

Listener

The load balancer endpoint that accepts traffic on a protocol and port, such as HTTPS on port 443.

Backend

A server instance that can receive forwarded work from the load balancer.

Backend pool

A named or logical group of backends eligible for a particular class of traffic.

Routing rule

A condition and action that decide which backend pool receives traffic.

Health check

A periodic probe used to decide whether a backend should remain eligible for new traffic.

Failover

Continued service after a component fails, using healthy capacity that remains available.

Active connection

A connection currently open on a backend. Duration determines how long it contributes to current work.

Traffic weight

A relative value that controls how much traffic one backend receives compared with others.

Upstream

A common proxy term for the backend service or pool to which traffic is forwarded.

The forwarding path

A valid path has three stages: traffic arrives, the load balancer makes a routing decision, and a healthy backend handles the work. A server drawn on the canvas but not connected to the load balancer contributes no usable capacity.

Clients
Incoming connections and requests
Load balancer
Selects an eligible backend
Backend pool
Healthy application instances
Figure 2. A load balancer is a traffic decision point. It accepts work, selects an eligible backend, and forwards the connection or request.
Common misconception

“A load balancer is what gives a system more capacity.”

It doesn't. A load balancer distributes traffic among instances that already exist. An autoscaler changes the number of instances. Many systems use both, but they solve different problems — a load balancer with three idle backends behind it has exactly as much capacity as those three backends provide, no more.

03
Layer 4 vs. Layer 7 load balancing

Choose What the Load Balancer Needs to Understand

Layer 4 and Layer 7 refer to layers in the network model. The distinction is entirely mechanical: a Layer 4 load balancer forwards a transport connection using addresses and ports and never opens it, so it cannot see anything about the HTTP request inside. A Layer 7 load balancer terminates the connection and reads the actual HTTP request — method, path, headers, cookies — before it decides where to send it.

Layer 4 load balancer

Forwards the TCP/UDP connection. Never opens the request itself.

Source and destination IP
Destination port, e.g. :443
Transport connection state
URL path — the request is unopened
Headers, cookies, request body
TCP :443
Connection metadata
Backend
Picked with no HTTP inspection
Layer 7 load balancer

Terminates the connection and reads the actual HTTP request.

Everything Layer 4 sees, plus:
HTTP method and URL path
Headers and cookies
Request body, if it chooses to inspect it
GET /media
Application request
Media pool
Route-specific backend
Figure 3. Layer 4 chooses a backend from a connection it never opens. Layer 7 opens the connection first — which is also why it costs more CPU per request.
QuestionLayer 4Layer 7
What can it inspect?IP, port, TCP or UDP connectionHTTP host, path, method, headers
Can it route /api and /media differently?NoYes
Typical useGeneral TCP services, very high throughputHTTP APIs, host and path routing, request policies
04
Mechanics

TLS Termination & the Reverse Proxy Role

A load balancer sitting in front of HTTPS traffic almost always plays a second role: it terminates TLS. The client's encrypted connection ends at the load balancer, not at the backend. The load balancer holds the TLS certificate and private key, decrypts the incoming request, and only then can a Layer 7 routing rule actually read the path, headers, or cookies described in the previous chapter — encrypted bytes can't be inspected for a path match.

Traffic between the load balancer and the backend is commonly plain HTTP again, since it stays inside a private network. Some deployments re-encrypt that hop for defense in depth, but either way the certificate lives in exactly one place — the load balancer — instead of being copied onto every backend instance.

Client
Encrypted HTTPS connection
Load balancer
Holds the certificate, decrypts the request
Backend
Receives a readable HTTP request
Figure 4. TLS termination ends the encrypted connection at the load balancer. Only after decryption can a Layer 7 rule actually read the path, headers, or cookies.
Why keep the certificate in one place?
A certificate copied onto every backend instance is a certificate that has to be rotated, secured, and audited on every backend instance. Centralizing TLS termination at the load balancer means renewing one certificate updates the whole fleet at once, and it removes the encryption/decryption CPU cost from every individual backend.
05
Load balancing algorithms

Choose the Signal That Represents Backend Work

A routing algorithm is the rule used to select a backend. The correct algorithm depends on whether servers are equal, whether request cost varies, and whether current connection count is a useful signal.

Round Robin
ABCABC

Equal request count. Best for similar servers and similar request cost.

Weighted Round Robin
224380

Configured traffic ratio. Best when backend capacities differ predictably.

Least Connections
7
2
5

Current open connection count. Best when connection durations differ.

Figure 5. There is no universally best algorithm. Choose the signal that best approximates current backend work for the workload in front of you.
Round Robin

Cycles through eligible backends in turn. It assumes each request costs roughly the same and each backend has roughly the same capacity.

Weighted Round Robin

Cycles according to configured weights. Deterministic, and appropriate when backend capacities differ in known proportions.

Least Connections

Chooses the backend with the fewest active connections. It reacts to long-lived requests that remain open while short requests finish.

Common misconception

“A load balancer just splits traffic evenly, no matter what.”

Round Robin does — but it's a deliberate default for a uniform fleet, not the definition of load balancing itself. Weighted Round Robin and Least Connections exist specifically because backends are often not identical and requests are often not equal-cost. Picking “even” when the fleet is heterogeneous overloads the smaller or slower-clearing backends while the rest sit under-used.

Check your understanding

A fleet has one Large backend and three Small backends, where Large has 4x the rated capacity of a Small. Which algorithm keeps every backend at roughly the same percentage of its own capacity?

06
Weighted load balancing

Route in Proportion to Backend Capacity

Round Robin sends equal request counts. That is unsafe for a mixed fleet because a Small server and a Large server receive the same traffic despite having different limits. Weighted Round Robin lets the routing share mirror capacity.

BackendRelative capacity and weightAssigned
Small
220 RPS
22
176 RPS
80%
Medium
430 RPS
43
344 RPS
80%
Large
800 RPS
80
640 RPS
80%
Large
800 RPS
80
640 RPS
80%
Figure 6. Weights 22:43:80:80 mirror capacities 220:430:800:800. At 1,800 RPS, every backend receives 80% of its own rated capacity.
Capacity-proportional weight
server weight ∝ server rated capacity
220:430:800:800 becomes 22:43:80:80
Assigned traffic
server RPS = total RPS × server weight ÷ total weight
1,800 × 22 ÷ 225 = 176 RPS for Small
Common misconception

“Load-balancer weights have to add up to 100, like percentages.”

Weights 22, 43, 80, and 80 sum to 225 in the figure above. The load balancer normalizes them into traffic shares automatically — they only need to represent relative capacity, not a literal percentage breakdown.

07
Health checks and failover

Stop Routing Before Spare Capacity Can Help

A failed backend must be removed from the eligible pool. Health checks perform this control-loop function by probing each backend and changing its routing state when failures cross a threshold.

Healthy state
Probe succeeds
Backend remains eligible
Traffic continues
Requests may be assigned
Unhealthy state
Probe fails
Failure threshold reached
Removed
Zero new requests sent
Figure 7. Health checks control eligibility. Spare capacity does not help if the load balancer continues sending requests to a failed backend.
Common misconception

“Health checks alone guarantee zero failed requests during a backend failure.”

There is always a detection window — the time between a backend actually failing and the health check crossing its failure threshold and removing it from rotation. Any request routed to that backend during the window still fails. Faster checks shrink the window; they cannot eliminate it.

Health detection and failover capacity are separate

Health checks prevent requests from reaching a failed server. They do not create replacement capacity. After removal, the remaining servers must still handle the specified failure traffic without crossing their utilization limit.

Normal operation
430
430
430
430
4 × 430 = 1,720 RPS
1,400 RPS demand · 81.4% utilization
One backend removed
430
430
430
430
3 × 430 = 1,290 RPS
1,000 RPS demand · 77.5% utilization
Figure 8. Failover capacity is calculated after removing the failed server. Normal capacity alone does not prove the service survives a failure.
Remaining capacity
normal fleet capacity − failed server capacity
1,720 − 430 = 1,290 RPS
Failover utilization
failure traffic ÷ remaining capacity × 100
1,000 ÷ 1,290 × 100 = 77.5%

What makes a useful health check

Representative

Probe a path that proves the service can perform meaningful work, not only that a process exists.

Cheap

The probe runs frequently, so it must not create significant application or database load.

Fast enough

Detection must be quick enough for the recovery objective without reacting to one harmless transient error.

Independent

Avoid a probe that declares every backend unhealthy because one optional shared dependency is degraded.

08
Removing a backend safely

Graceful Deregistration & Connection Draining

Health checks answer whether a backend is eligible for new traffic. They don't answer what happens to the requests already in flight on that backend the moment it's marked unhealthy, or simply scheduled for planned removal — a deploy, a scale-down, a maintenance window.

A hard removal stops sending requests instantly and drops any connection already in progress mid-response. Connection draining (also called graceful deregistration) stops sending new requests immediately, but lets requests already in flight finish normally, up to a configured drain timeout. Only once every in-flight request completes — or the timeout expires — is the backend fully removed.

Hard removal
New request — stopped
In-flight request — dropped mid-response

The backend is removed instantly. Anything already running on it is cut off.

Connection draining
New request — stopped immediately
In-flight request — allowed to finish

New traffic stops right away, but existing requests get to complete before removal finalizes.

Figure 9. Both paths stop new traffic instantly. Only draining protects the requests that were already in progress.
Drain timeout heuristic
drain timeout ≥ slowest expected request duration
Too short and draining degrades into a hard removal anyway
Why does this matter for deploys, not just failures?
Every rolling deploy removes a healthy backend on purpose. Without draining, every deploy silently drops whatever requests happened to be mid-flight on that instance at that exact moment — a self-inflicted version of the same failure health checks exist to catch.
09
Connection-aware routing

Request Count Is Not Current Work

Round Robin equalizes arrivals, not active work. When all requests finish in similar time, that distinction may not matter. When some requests remain open far longer, one backend can accumulate more active work than another even after receiving the same number of requests.

Fast API calls50 ms

Many requests complete quickly, so each connection disappears almost immediately.

Data exports8 seconds

Only a few exports arrive each second, but they remain open long enough to dominate active connection count.

Figure 10. Request rate measures arrivals. Active connections also depend on duration, so equal request counts can still create unequal current work.
Fast-request concurrency
1,800 RPS × 0.05 s = 90 active requests
Export concurrency
15 RPS × 8 s = 120 active exports
Total average concurrency
90 + 120 = 210 open requests
Only 15 of 1,815 arrivals per second are exports, yet exports account for 120 of 210 active requests.
Key idea

Round Robin distributes arrivals evenly. It cannot distribute active work evenly, because arrivals and active work are only the same thing when every request takes the same amount of time.

Least Connections uses open connection count as a live approximation of backend work. It does not replace capacity planning. The fleet still needs enough RPS capacity, and connection count is only useful when it correlates with resource usage.

10
Path-based routing and backend pools

Route Each Workload to Dedicated Capacity

A Layer 7 routing rule can inspect the URL path and choose a backend pool. This allows one domain to serve multiple workloads while keeping their compute capacity independent.

Layer 7 path rules
/api
2 × Large
1,200 RPS
/media
1 × Large
600 RPS
/admin
1 × Medium
200 RPS
Figure 11. Path-based routing maps each URL path to a dedicated backend pool. Pool isolation prevents one route from consuming another route's compute budget.

Pool isolation

A backend is dedicated only when it belongs to one pool. Connecting the same application server to two route pools creates shared capacity and allows one route to affect the other. Isolation requires separate membership, separate sizing, and separate utilization checks.

/api capacity
1,200 ÷ 0.75 = 1,600 RPS
2 × Large
/media capacity
600 ÷ 0.75 = 800 RPS
1 × Large
/admin capacity
200 ÷ 0.75 = 266.7 RPS
1 × Medium
Why does the practice canvas show a separate load-balancer node per path?
The section exercise represents each path rule as a separate Layer 7 load-balancer node because canvas nodes don't carry editable route labels. In production, one Layer 7 proxy commonly owns several path rules that select several backend pools.
11
Routing design workflow

A Repeatable Interview Method

Separate topology, algorithm, health, and capacity decisions. A correct answer should explain what traffic is being routed, which signal drives selection, which backends are eligible, and what happens when one fails.

  1. 1

    Identify the traffic unit

    State whether the input is requests, connections, jobs, bytes, or sessions. Keep units attached to every calculation.

  2. 2

    Choose Layer 4 or Layer 7

    Use Layer 7 when routing depends on HTTP host, path, method, or headers. Use Layer 4 when transport-level forwarding is sufficient.

  3. 3

    Decide where TLS terminates

    If the routing rule needs to read the request, TLS has to terminate before that decision is made.

  4. 4

    Define backend pools

    List which servers may handle each traffic class. Decide whether capacity is shared or isolated.

  5. 5

    Choose the routing algorithm

    Use Round Robin for a uniform fleet, Weighted Round Robin for known capacity differences, and Least Connections for varying connection duration.

  6. 6

    Calculate normal capacity

    Verify aggregate throughput and every backend's utilization under the selected traffic split.

  7. 7

    Define health and eligibility

    Specify how failed backends are detected and removed from rotation, and how in-flight requests are drained on planned removals.

  8. 8

    Calculate failure capacity

    Remove the failed backend from the math and verify that the remaining fleet serves the required failover load.

  9. 9

    Check isolation and cost

    Verify that no backend is unintentionally shared and that every load balancer and server fits the monthly budget.

Do not trust aggregate capacity

A bad traffic split can overload one backend while total fleet capacity looks sufficient.

Do not ignore duration

Equal arrival counts do not imply equal active work when requests have different lifetimes.

Do not fake isolation

A backend shared by two pools is still a shared failure and capacity domain.

12
Practice labs

Apply Each Traffic-Routing Concept

Each lab isolates one routing decision. Build mode asks you to construct the system from a specification. Fix mode gives you a concrete routing failure and asks for the smallest effective repair.

01

Uneven Fleet

Use Weighted Round Robin to distribute traffic across backends with different capacities.

02

The Dead Server Still Gets Traffic

Enable health checks and prove the remaining fleet can carry failover traffic.

03

Long Requests Break Round Robin

Use Least Connections when active connection duration matters more than arrival count.

04

Route by Workload

Use Layer 7 path routing and dedicated backend pools for independent workloads.

Ready to route traffic

You are ready when you can justify the routing layer, algorithm, health behavior, backend membership, failure capacity, and cost without relying on the diagram alone.

Start section