The Engineer’s Field Guide to Load Balancers Load balancers look deceptively simple “split traffic across servers”, but they’re really the circulatory system of modern systems. They terminate connections, speak multiple protocols, apply routing policies, enforce security controls, measure health, and hide failure, all while keeping latency in the tens of milliseconds (or less) and throughput in the millions of requests per second. This guide goes deep: how load balancers actually work, the trade-offs behind common algorithms, what to tune What a Load Balancer Is (and Isn’t) At heart, a load balancer (LB) is a programmable network intermediary. As a reverse proxy, it accepts client traffic on a virtual IP (VIP), chooses a backend (“upstream”, “origin”, “target”), and forwards the traffic. Depending on where it operates in the stack: [object Object], [object Object] There are a few deployment archetypes: [object Object], [object Object], [object Object] An LB is not automatically a WAF, API gateway, or service mesh; but modern L7 LBs overlap with those roles. The mental model that scales is control plane vs. data plane: the data plane is the hot loop that accepts, routes, and forwards packets; the control plane distributes configuration and membership (“these are the healthy instances, with these weights”). The Data Path: Step by Step When a request hits a reverse proxy, roughly this happens: [object Object], [object Object], [object Object], [object Object], [object Object], [object Object], [object Object] Policies & Algorithms (and When They Bite) Choosing where to send traffic seems simple until tail latency shows up. Here’s what really matters. Round Robin / Weighted Round Robin. Evenly distributes requests; with weights, bigger machines do more. It ignores in-flight work and can overload a slow instance. Least Connections (or Least Requests). Prefers backends with fewer active requests. Better under heterogenous latency, but can be gamed by long-lived connections (e.g., WebSockets). Many implementations use EWMA of observed latency to approximate “least loaded”. Random with “Power of Two Choices”. Sample two backends at random and pick the less loaded. With negligible overhead, this dramatically reduces worst-case queueing. Consistent Hashing. Hash a key (user ID, session, cache key) to pick a backend; when membership changes, only a small fraction remaps. Variants: Rendezvous/HRW, Jump hash, Maglev. Great for caches and sticky state, risky if a shard gets hot. Sticky/Affinity. Keep a client on the same backend (via cookie, source IP, TLS session ID, or QUIC CID). Essential for stateful apps, but it undermines elasticity: a hot client can scorch a single node. Prefer stateless sessions or server-side stores when you can. Slow-start Warmup. After adding a backend, ramp traffic gradually while caches warm and JIT/GC settles. Helps avoid “cold start” spikes. Outlier Detection / Ejection. Temporarily remove backends that exceed error or latency thresholds. Combine with passive health checks (mark on 5xx/timeout) and active checks (HTTP/TCP probes). Connection Management: The Performance Bedrock Keep-alives & pools. Reusing connections to backends shaves RTTs and TLS handshakes. Size pools conservatively; too many idle connections starve ephemeral ports and memory. Idle timeouts. Clients, LBs, and servers each have their own. Mismatches cause mysterious disconnects (especially with gRPC or WebSockets). Align them and use heartbeats/pings. TCP specifics. [object Object], [object Object], [object Object], [object Object] HTTP/2 and HTTP/3. HTTP/2 multiplexes streams over one TCP connection (watch for server-side head-of-line at LB↔backend if only one connection is used). HTTP/3/QUIC runs over UDP; L4 LBs must use CID-aware routing for affinity because 5-tuple changes across NATs and migrations. WebSockets & long-lived streams. Treat separately: longer idle timers, periodic pings, and connection draining logic that waits for streams to finish. TLS at the Edge (or Not) Terminate at the LB to centralize ciphers, certificates, and HSTS/OCSP stapling. Then either: [object Object], [object Object] Practical details: [object Object], [object Object], [object Object], [object Object] Health Checking That Tells the Truth Active checks (HTTP/HTTPS/TCP) run on intervals with timeouts and “N-of-M” thresholds. Jitter intervals to avoid herding. Check the thing the user needs: for HTTP, hit /healthz that probes dependencies lightly; for gRPC use health-check protocol; for TCP ensure banner/readiness. Passive checks demote backends on real traffic errors/timeouts. Combine with outlier detection (500s, resets, slow EWMA) to protect users during partial failures. Drain & graceful shutdown. On deploy/scale-down, mark instances as draining so new requests go elsewhere while existing ones finish. Pair with slow-start when they come back. Reliability and High Availability No single LB. Use at least two per zone; for on-prem, float