Generate a follow-up sub-lesson on any aspect of this topic
Guide complete
Latency, Throughput, And Bandwidth Intuition
Generate a follow-up sub-lesson on any aspect of this topic
Latency, Throughput, And Bandwidth Intuition
TLDR;
- A feed can feel slow at 300 ms even when the backend is barely busy, because users notice waiting, not utilization.
- You can hit 200 req/s at 300 ms average latency only if enough requests are in flight at once.
- Batching, caching, and compression often trade lower bandwidth for higher tail latency when queues form.
A photo sharing feed is a good place to build intuition because one screen load mixes database work, server compute, and network transfer in a single, user felt wait. Suppose the app is simple and we start with one app server talking to one database, and the mobile client does one GET /feed per screen load. Users report the feed feels slow when that screen takes about 300 ms end to end, even though 300 ms sounds small in isolation. The reason is that this is wall clock waiting time, and it includes every pause the user cannot parallelize away. Let’s ground that baseline request path in a simple picture.
The diagram has four moving parts, the mobile client, the internet path, one app server, and one database, and one number that matters to the user, the 300 ms from tap to feed rendered. That 300 ms is not a single thing. It is the sum of multiple smaller waits, some on CPU, some on the network, and some inside the database engine, and the user experiences them as one delay. In the rest of this guide we keep this same feed request and change constraints around it, so the tradeoffs stay concrete.
Latency is the time to complete one operation from start to finish, like one GET /feed taking 300 ms. Throughput is how many operations a system completes per unit time, like 200 feed requests per second across all clients. Bandwidth is how many bytes per second can move across a link, like a 20 MB/s cellular downlink or a 1 Gbps NIC on the app server. All three apply to the same feed load, and confusing them leads to wrong fixes, like adding servers when the problem is oversized responses on a narrow link. The mapping is easier to remember when we anchor it to a single mental model.
In the mapping, the feed request’s latency corresponds to travel time for one car, throughput corresponds to how many cars per second pass a checkpoint, and bandwidth corresponds to how much cargo per second the road can carry. A key consequence is that increasing bandwidth can reduce latency only when the response is large enough that transfer time dominates, and increasing throughput can leave latency unchanged when each request still has the same number of round trips and the same per request work. That is why a system can feel slow to one user while still serving many users per second.
For a single feed load, latency is the sum of several stages that often happen serially. The client sends the request, the load balancer forwards it, the app server authenticates and builds a query, the database reads indexes and rows, and then the response bytes move back to the client where the app parses and renders. Even when compute is fast, network round trips add fixed waiting that does not improve just because there are more servers, since each request still must cross the network. Adding more app servers mainly helps when the app server is the bottleneck for many concurrent users, raising throughput by sharing load. The per request latency only drops if the critical path gets shorter, meaning fewer round trips, less waiting on locks, faster queries, or fewer bytes transferred.
The sequence shows one client request flowing through the load balancer to an app server and then to the database, with network hops that create round trip delays even when each component is lightly loaded. The important detail is that the app server and database can be fast in isolation, but the request still waits for network acknowledgments and for the database to return results before the app server can serialize a response. Scaling out app servers increases the number of requests that can be handled at once, but it does not remove those round trips, so it often increases throughput without materially lowering latency.
Design rule
When one request must wait on two round trips, extra servers do not erase the waiting.
Two back of the envelope relationships keep latency, throughput, and concurrency from feeling abstract. Little’s Law is a steady state relationship that says the average number of in flight items equals throughput times average time in system , written as . In our feed scenario, if we want requests per second and each request takes seconds on average, then the system must sustain in flight feed requests at once. That does not mean 60 threads. It means some combination of concurrent connections, async work, and database queries active at the same time, without queues exploding. The second relationship is the bandwidth delay product (BDP), which is the amount of data a link can have in flight, roughly bandwidth times round trip time, and it explains why high latency links need more in flight bytes to fully utilize bandwidth.
With the calculator set to 200 req/s and 300 ms, the output should show about 60 in flight requests needed to meet that throughput at that latency. If the system only allows 20 concurrent in flight requests before queuing, throughput cannot reach 200 req/s without increasing latency, because the extra 40 requests per second pile up in a queue and wait. This is the moment where throughput targets turn into explicit concurrency limits and backpressure behavior, instead of wishful scaling assumptions.
Latency numbers become useful when you treat them as orders of magnitude rather than precise constants. CPU cache hits are measured in nanoseconds, RAM in tens of nanoseconds, SSD reads in tens to hundreds of microseconds, HDD reads in milliseconds, and network hops climb from sub millisecond in a zone to tens of milliseconds across regions. The practical move is to ask which tier you are paying, because crossing a tier boundary is often a 10x to 1000x jump. In the feed request, a database query that spills to disk or a cache miss that forces a cross region fetch is not a small regression, it is a new dominant term in the critical path.
The table highlights the scale gaps, like cache versus RAM, storage versus memory, and same region versus cross region network. The takeaway is not the exact value of any row, it is which categories belong to microseconds and which belong to milliseconds. When the feed’s 300 ms budget is tight, one unexpected millisecond class dependency can consume the entire margin, and the fix is usually architectural, like caching closer, not micro optimizing code.
Most optimizations move one of latency, throughput, or bandwidth by pushing cost into another dimension, and the failure mode is usually queuing or tail latency. The common mechanisms show up in feeds because feeds have lots of repeat reads, variable payload sizes, and bursty traffic.
The most important interaction is saturation. When CPU, database connections, or bandwidth approach 100 percent utilization, small bursts create queues, and then latency rises sharply while throughput flattens. That is why raising throughput targets often requires explicit concurrency limits and admission control, not just more workers, so the system fails by shedding load instead of dragging every request into multi second waits. Now we can make those interactions tangible by adjusting batching and parallelism.
After tuning, notice which metric bends first. Larger batches often increase throughput and reduce bytes per item, but they also push up p95 latency because some requests wait for the batch window. More parallelism often reduces average latency until the database or network saturates, then p95 climbs as queues form. The knob that feels like free performance is usually just moving time from transfer to compute, or from average latency to tail latency.
Practical caveat
If p95 rises while average stays flat, a queue formed somewhere, even if graphs look calm.
For the photo feed, three decisions usually pay off before any exotic work. Reduce round trips in the critical path, cap payload size so the network transfer term stays bounded, and set concurrency limits so the system sheds load before queues explode. Reducing round trips can mean consolidating database reads, denormalizing for reads, or using a cache that returns most feed items without a database hit. Capping payload size can mean pagination, smaller images by default, and avoiding unbounded metadata. Concurrency limits can mean a max number of in flight requests per app server and a max number of database connections, so stays near the designed steady state rather than turning into an unbounded queue. The right choice depends on whether the dominant cost is waiting on the network, waiting on the database, or moving too many bytes.
In the comparison, latency oriented work tends to remove waits like extra round trips, throughput oriented work tends to add capacity like more replicas or more workers, and bandwidth oriented work tends to shrink responses through smaller images or compression. For our 300 ms feed, a redesign trigger is when cross region traffic becomes normal, because the fixed network component jumps, or when p95 exceeds 1 second, because that usually signals saturation and queuing rather than slow code. At that point, the system needs a new shape, like regionalization, read replicas closer to users, or a more cache forward design, because the original single region request path cannot meet the same user felt constraint anymore.