Generate a follow-up sub-lesson on any aspect of this topic
Guide complete
Choosing TCP vs UDP in System Design
Generate a follow-up sub-lesson on any aspect of this topic
Choosing TCP vs UDP in System Design
TLDR;
- A single TCP connection can make audio sound late even when bandwidth looks fine.
- UDP can keep live audio responsive, but only if the app handles drops and reordering.
- QUIC often wins when you need UDP like latency with stream level reliability options.
A chat and live audio app makes the tradeoff concrete because the system needs two different kinds of correctness at the same time. Chat messages need to arrive and read in order, while audio needs to stay under a tight end to end latency budget so people do not talk over each other. Start with a baseline that keeps complexity low by running both chat and audio over one client to server TCP connection in a single region, and assume an RTT of 150 ms between a mobile client and the server edge. Here is that baseline transport view at a high level.
The key observation in the baseline is that one transport behavior applies to both workloads, even though the user experience expectations are different. The load balancer and app server see a single byte stream, so any delay that affects one part of the stream can indirectly affect the other. This is why protocol choice in design discussions usually starts by separating what must be correct from what must be timely.
Transport choice is usually decided by four measurable signals that show up as user visible symptoms in chat plus live audio. Round trip time (RTT) is the time for a packet to go to the server and for the response to come back, so it sets the scale of handshake time and any recovery that needs a round trip. Packet loss is the fraction of packets that never arrive, which forces either waiting for recovery or skipping data. Jitter is variation in one way delay, which forces buffering if you want smooth playout. Bandwidth is the sustained rate you can send, which bounds bitrate and how much redundancy you can afford. The mapping from metrics to symptoms is easiest to see when we line them up directly.
When RTT is 150 ms, any mechanism that needs one extra round trip adds a noticeable chunk of time to the audio budget, even if it only happens occasionally. When loss or jitter increases, audio either gets gaps or it gets buffered, and buffering raises latency. Chat is more forgiving about a 300 ms wait, but it is not forgiving about missing or reordered text, which is why the same metrics can argue for different transports inside the same product.
TCP is a good default when the system values complete, ordered delivery because it turns an unreliable network into a reliable in order byte stream. Three way handshake sets up the connection using SYN, SYN-ACK, and ACK, then the sender transmits bytes that the receiver acknowledges with acknowledgments (ACKs). If a segment is missing, TCP uses retransmission to send it again, and it only delivers data to the application in order, which means later bytes wait behind a gap. The moment loss happens, recovery time is tied to RTT because the sender needs feedback before it can confidently repair the stream. The control flow looks like this for one lost segment at 150 ms RTT.
In the sequence, the client and server spend about one RTT establishing state before any application data moves, then a lost data segment forces at least one additional RTT worth of waiting before the missing bytes can be repaired and delivered in order. The important part is not the exact TCP algorithm, it is the shape of the penalty. TCP converts loss into delay, and the size of that delay is dominated by RTT and the retransmission timeout behavior.
Design rule
If an interaction cannot tolerate waiting an extra RTT, avoid depending on retransmission.
TCP reliability is expensive for real time media because the audio pipeline cares more about playout time than perfect completeness. Head of line blocking is when one lost TCP segment prevents later segments from being delivered to the application, even if the later segments arrived. For live audio, those later bytes might contain newer audio frames that would have been useful, but TCP holds them until the missing data is repaired. TCP also runs congestion control, which reduces send rate after loss to avoid overloading the network, and that backoff can create bursts and gaps that look like jitter to the audio playout logic. When loss and jitter move around, TCP tends to push the app toward larger buffers, and buffers raise end to end latency. Let’s dial loss and jitter and watch what happens relative to a 200 ms audio latency target.
As the network gets slightly worse, TCP spends more time waiting for in order delivery and more time adapting its send rate, so the app compensates by buffering more audio before playing it. The result is that the audio sounds continuous but arrives late, which is worse than the occasional small glitch in a conversational product. This is the key tension. TCP is keeping its promise, but the promise is not aligned with the audio user experience.
UDP is a good default when the system values getting the freshest data to the receiver quickly, even if some data is missing. A UDP sender emits independent datagrams, which are message like packets that are not part of a reliable byte stream. There is no handshake at the transport layer, no built in retries, and no requirement that datagrams arrive in order, or arrive at all. Best effort means the network will try to deliver datagrams, but drops, duplication, and reordering are all normal outcomes that the application must be prepared to observe. Here is what that looks like for audio frames when the network drops and reorders packets.
The key difference in the flow is that the receiver immediately sees whatever arrived, including newer frames that happened to beat older frames. That gives the application the chance to favor timeliness, for example by discarding late frames rather than waiting for them. UDP does not fix the network, it exposes the network, which is exactly what real time media systems often want.
Making UDP usable for audio usually means rebuilding a small, targeted subset of reliability at the application layer. Sequence numbers let the receiver detect reordering and loss without waiting for transport level repair. A jitter buffer is a small queue that delays playout by a controlled amount so the receiver can smooth out variable arrival times. Forward error correction (FEC) adds redundant parity so the receiver can reconstruct some losses without a round trip. Selective retransmission can still exist, but it is usually reserved for data where being late is better than being missing, such as a video keyframe or an occasional control message, not for every audio frame. Here are the building blocks and how each one pushes on latency.
The typical shape is that sequencing is almost free, jitter buffering costs a fixed amount of added delay, and FEC costs bandwidth instead of time. That trade lets the system stay under a conversational threshold by spending bytes to avoid waiting. The app ends up with knobs that TCP does not give you, but it also takes on complexity and more failure modes, such as choosing a jitter buffer that is too small for current jitter or too large for the latency budget.
Practical caveat
A jitter buffer that hides jitter also hides problems until users complain about delay.
In design discussions, protocol choice lands faster when it is anchored to what correctness means for the workload. TCP is usually the right pick for APIs, file transfer, and many database client connections because application correctness depends on complete ordered delivery and the cost of occasional extra RTTs is acceptable. UDP is usually the right pick for VoIP, gaming, and live video because the newest state is more valuable than perfect delivery of old state. Quick UDP Internet Connections (QUIC) is a transport built on UDP that adds encryption and multiplexed streams, reducing connection setup cost and avoiding TCP head of line blocking across streams, while still running over UDP for deployment practicality. The trade space across TCP, UDP, and QUIC looks like this.
The matrix typically shows TCP as simpler for the application but more likely to convert loss into delay, UDP as lowest setup overhead but highest app responsibility, and QUIC as a middle ground that can carry multiple independent streams so one stream’s loss does not block the others. For the chat plus audio scenario, that often suggests separating chat control traffic from audio media traffic, even if both ultimately traverse the same load balancer and edge.
For the chat and live audio app, splitting transports aligns the protocol promise with the user experience. Put chat and control on TCP, or on a reliable QUIC stream, so messages arrive completely and in order. Put audio on UDP, or on QUIC datagrams when available, so the receiver can keep playout close to real time using sequencing, a small jitter buffer, and optional FEC. The split also limits blast radius. A burst of audio loss does not stall chat delivery, and a chat retransmission spike does not drag audio latency upward. Two constraint changes commonly flip the choice, and we can reveal them explicitly.
If the environment includes strict corporate firewalls or NAT behavior that blocks UDP, the system often moves toward TCP or QUIC streams because reachability becomes the primary constraint. If the product requirement changes to perfect, record quality audio where every sample must be preserved, then the system shifts toward a reliable transport behavior, or adds heavier recovery on top of UDP, and accepts higher latency or post processing. The durable design decision is not TCP versus UDP in isolation, it is whether the workload prefers waiting to be correct or skipping to stay timely, and what external constraints force that preference to change.