Generate a follow-up sub-lesson on any aspect of this topic
Guide complete
Killing Single Points Of Failure
Generate a follow-up sub-lesson on any aspect of this topic
Killing Single Points Of Failure
TLDR;
- A single checkout dependency can turn 500 orders per minute into zero with no partial fallback.
- Stateless redundancy removes the fastest outages, but it does not protect writes or customer carts.
- Your final design is defined by what you still let be single, and what would force multi-region.
A single point of failure is any one component whose failure stops the whole user-critical path, even if everything else is healthy. In an online store checkout, the simplest architecture has exactly one load balancer, one app server, and one primary database, so any one of those boxes dying drops revenue to zero instead of degrading gracefully. Let’s anchor on one workload so the tradeoffs stay concrete, 500 orders per minute at a 200 dollar average order value, which is dollars per minute of checkout throughput riding on three singletons. Here is that baseline checkout path at a glance.
The important thing in the diagram is not the shapes, it is the fact that every request has only one valid next hop at each step, so there is no alternate route when the load balancer, app host, or database host fails. Under a fail-stop assumption, which means a component halts and stops responding rather than returning incorrect answers, the checkout API typically times out, retries briefly, then surfaces an error to the customer, and the system cannot earn its way out of the outage.
A practical SPOF scan treats the architecture like a dependency graph and asks one repeatable question at every hop, what happens if this box dies right now. The scan is most useful when you do it from the user action inward, starting at the client request, then the load balancer, then the app, then the database, because that order mirrors how failures present and how blast radius propagates. The output of the scan is a list of components that turn a partial failure into a full checkout outage, plus the first mitigation that would let the path keep working. Here is a structured way to record that per component.
The key observation is that a component is a SPOF when all its downstream dependencies are reachable only through it, so its failure makes the checkout path unreachable rather than merely slower. The widget’s mitigation column also forces a distinction between redundancy, which provides another instance of the same tier, and replication, which provides another copy of state, because the fix depends on whether the tier is stateless or stateful.
Redundancy works best first on stateless tiers because you can add instances without needing to keep user data consistent between them. For checkout, that usually means at least two load balancers or an equivalent highly available entry tier, and at least two app instances behind it, with health checks that stop routing to a dead instance before customer requests pile up.
The before and after topology makes the “one and only one” routing path visible.
In the after view, checkout requests still follow one logical path, but the load balancer selects among multiple healthy app instances, and the apps can restart or be replaced without making the whole service unavailable. This removes the fastest and most common SPOFs, but it does not yet protect the state that makes a checkout a checkout, inventory reservations, payments, and order records still live in one database.
Design rule
Stateless tiers fail over by routing, stateful tiers fail over by authority.
Replication addresses the SPOF in the stateful tier by keeping at least one additional database node with a copy of the data that can take over if the primary fails. A replica is a database instance that continuously applies changes from a primary so it can serve reads, and sometimes be promoted to accept writes, but the failure behavior depends heavily on whether replication is synchronous or asynchronous.
With synchronous replication, a write is acknowledged only after the replica confirms it has the change, which reduces data loss but can reduce write availability when the replica is slow or unreachable. With asynchronous replication, the primary acknowledges writes immediately and ships changes after, which preserves write availability during replica slowness but creates replication lag, and a primary failure can lose the last unreplicated orders. The tradeoff is easiest to feel when we map it onto the 500 orders per minute workload.
The tuning view is really about where you want to pay, in steady state latency and fragility with synchronous replication, or in potential recovery data loss with asynchronous replication. In both cases, a primary failure still takes writes down until something becomes the new write leader, so replication is necessary but not sufficient without a promotion path.
Failover is not just starting the replica, it is safely changing which node is allowed to accept writes while preventing two nodes from believing they are the leader. A leader is the single database node that is currently authoritative for writes, and changing leaders requires detection, decision, and client reconnection that converge quickly enough to meet an RTO target.
The mechanics usually include timeouts for failure detection, a leader election or an external coordinator to decide the new leader, and fencing that prevents the old leader from accepting writes if it comes back with stale authority. Fencing can be a lease, which is a time-bounded right to be leader, or a mechanism that cuts the old primary off from shared storage or the network so it cannot keep writing. This is the minimal sequence of events you are trying to make boring.
The sequence matters because each step is a place you can accidentally extend downtime, for example when detection timeouts are too long, or cause inconsistency, for example when split brain lets two primaries accept writes. If clients reconnect only after a promotion completes, you trade a short hard error window for a consistent database, which is typically the right trade for checkout because duplicate or lost orders are often worse than a minute of unavailability.
The real decision between active-active and active-passive is how much recovery speed you buy with ongoing operational complexity and consistency risk. RTO is the maximum acceptable time to restore service after a failure, and RPO is the maximum acceptable data loss measured in time, which together force clarity on whether you are optimizing for availability, durability, or simplicity.
Active-passive usually means one leader taking writes and one standby ready to be promoted, which keeps the write authority simple but makes RTO depend on detection and promotion. Active-active usually means two sites or clusters both serving live traffic, which can deliver very low RTO but raises hard questions about conflict resolution, global consistency, and how to prevent correlated failures from taking both sides down. The comparison is easiest when pinned to explicit RTO and RPO targets and the failure modes you can actually operate.
The decision matrix should make one thing unavoidably clear, active-active is not a free upgrade from active-passive, it is a different system with different failure behaviors. If your checkout requires near zero RPO, active-active still does not guarantee it unless you also accept the latency and coupling of synchronous cross site replication, which can turn network hiccups into global write stalls.
After you add redundancy and replication, the remaining outages often come from shared dependencies that were invisible in the original sketch because you assumed they were always there. These are SPOFs that sit off to the side of the request path but still gate the entire system, for example a single configuration store that both app instances need at startup, a single DNS zone that controls where checkout traffic goes, or a single certificate authority process that can block renewals.
Starting from the improved diagram, we can uncover which dependencies are still shared.
The point of the reveal is blast radius, two app instances do not help if both need the same failing dependency to start, and two database nodes do not help if the same network partition isolates both from clients. This is also where you separate “can’t fail” dependencies from “can fail but degrade” dependencies, since checkout can sometimes degrade, for example by refusing new orders while still showing past order status.
A good anti-SPOF design ends with a list of remaining singletons you explicitly accept, because eliminating every last one often costs more complexity than it saves. For checkout in one region, it is defensible to accept some singletons that only affect operability, like a single metrics pipeline, while refusing singletons that directly gate routing, write authority, or certificate validity for customer traffic.
The trigger to redesign toward multi-region is not that a component is still single, it is that a single regional event can violate your RTO or RPO in a way the business cannot tolerate, like a cloud region outage that takes both the active database and its replicas with it. When that trigger is real, the design question shifts from “how do we add another instance” to “how do we bound correlated failure,” which is the core reason you kill single points of failure in the first place.