Generate a follow-up sub-lesson on any aspect of this topic
Guide complete
Load Balancing Fundamentals
Generate a follow-up sub-lesson on any aspect of this topic
Load Balancing Fundamentals
TLDR;
- A single server can handle req/s until one failure takes checkout fully down.
- L4 and L7 balancers fail in different ways because they can observe different protocol layers.
- Sticky sessions reduce logins breaking but can make failures and scaling harder to reason about.
A checkout web app can feel fine on one server right up until traffic and failure arrive at the same time. At requests per second, the single app server becomes both the throughput bottleneck and the single point of failure, since every user request depends on that one process and that one machine staying healthy. The database can still be a separate node, but it does not help if the only web tier node is saturated or rebooting. One useful baseline is to name the components and the blast radius of a crash. When the app server dies, the user impact is not partial degradation. It is total unavailability of the web tier until the node returns or is replaced, even if the database is fine. Here is the baseline shape we are starting from:
A load balancer exists to make many servers look like one reliable entry point while it spreads work across the pool. A common way to do that is a virtual IP (VIP), which is a stable address clients connect to while the balancer chooses a backend for each connection or request. The balancer also becomes the place where you can enforce timeouts, retry policy boundaries, and consistent routing rules, instead of duplicating those decisions in every client. This only removes the single point of failure if the balancer itself is not a single node. In practice that means running two balancers in active active mode behind a shared VIP or under DNS, so one can keep forwarding when the other fails. The key blocks and what breaks when one is missing are easiest to see in one diagrammed set of building blocks:
Layer 4 load balancing makes routing decisions using transport information like TCP or UDP tuples, which means it can be fast and protocol agnostic but blind to URLs, headers, and cookies. It typically forwards whole connections and cannot easily send /checkout to one pool and /images to another without additional hints, since it is not parsing HTTP. This simplicity can reduce overhead and failure surface, but it also limits observability and fine grained policy.
Layer 7 load balancing operates on application data like HTTP methods, paths, and headers, so it can do content based routing, header based canaries, and more explicit rate limiting. It can also terminate TLS, which means it decrypts and re encrypts traffic, improving control and metrics at the cost of extra CPU and a larger security and configuration footprint. The trade is not only features versus latency, but also what breaks when parsing, certificates, or header expectations go wrong. Compare the two layers side by side:
A balancing algorithm is the rule that decides which backend gets the next connection or request. Round robin is simple and works when backends are truly identical, while least connections adapts when some servers are already busy. Weighted variants matter as soon as servers are not equal, since sending equal traffic to a smaller instance can increase queueing and tail latency for everyone.
Some workloads are uneven not because servers differ, but because keys differ. Consistent hashing is a routing method that maps a key like session_id or user_id to a backend in a way that minimizes remapping when backends are added or removed, which can stabilize cache hit rates and reduce churn. It is a good fit when one endpoint is hot and you want per key stickiness without fully pinning every user to one server forever. Try the same pool under different rules and watch where the load concentrates:
Health checks are how the balancer decides which backends are eligible to receive traffic. An active check is a periodic probe like an HTTP GET /healthz or a TCP connect attempt, while passive detection uses real traffic failures like timeouts or connection resets as evidence that a backend is unhealthy. The balancer then needs thresholds and timers so it does not eject a server for one blip, and it needs a way to add it back gradually so a cold server does not take a full share immediately.
Most operational pain comes from flapping, where a backend repeatedly flips between healthy and unhealthy. Checks that are too aggressive for the actual failure mode cause churn, and churn looks like random latency spikes because the pool size keeps changing. Two knobs dominate this behavior. The check interval and timeout decide how quickly you notice failure, and the healthy and unhealthy thresholds decide how much evidence you require before changing state. Walk through a failure timeline and see how those knobs change removal and readd decisions:
Sticky sessions keep a user bound to the same backend across requests, usually by a cookie inserted by the balancer or by hashing the source IP. This helps when the app server holds session state in memory, since a plain round robin algorithm will send the next request to a different server that does not have that user session. The visible symptom is users being logged out, carts disappearing, or multi step checkout flows failing mid way. Affinity has costs that show up under failure and scaling. If one backend dies, every user stuck to it experiences a hard break until they are reassigned, and the reassignment can overload remaining nodes because the balancer is redistributing stateful load suddenly. A more durable alternative is to make the app stateless and move session state to a shared store like Redis, which makes any backend able to serve any request but adds a new dependency that must be scaled and made highly available. Compare the trade offs directly:
For the req/s checkout app, a practical default is an L7 balancer if you need HTTP aware routing, TLS termination, and clear per endpoint metrics, otherwise L4 is often simpler and easier to keep fast under load. With a mixed capacity pool, weighted least connections is a reasonable starting point because it reacts to transient imbalance while respecting that a larger instance should take more work. Health checks should be conservative enough to avoid flapping, with a real dependency check for the app such as database connectivity only if you are comfortable failing closed when the database is degraded. The constraint that usually forces a redesign is state. If the app remains stateful, sticky sessions become a crutch that makes failover events user visible, and the long term fix is pushing state into a replicated store that can survive node loss. A different forcing function is multi region traffic, since global load balancing turns into a CAP shaped choice between low latency local reads and consistent cross region checkout state. Load balancing is the tool that makes a server pool behave like one entry point, but the moment state crosses that boundary, the system stops being just routing and becomes replication and consistency.