Generate a follow-up sub-lesson on any aspect of this topic
Guide complete
How DNS Resolution Works
Generate a follow-up sub-lesson on any aspect of this topic
How DNS Resolution Works
TLDR;
- A single DNS server design collapses under multi region load and failure.
- One cold DNS cache can add multiple network round trips before any HTTP byte.
- TTL choices decide whether failover takes seconds or feels stuck for hours.
A web page load starts with turning a human name like www.acme-bakery.example into an IP address the network can route to, and that name to address lookup is what DNS does. A simplistic baseline is a hosts file entry or a single DNS server that returns one A record, which works until you have users across regions where latency and outages are normal. At that point one server becomes a hotspot, a single point of failure, and a single region answer becomes wrong for distant clients. Here is the baseline flow we are trying to outgrow.
The baseline diagram has three moving pieces, the browser or OS stub, one DNS server, and one returned address record, and it breaks because every client depends on the same server being reachable and fast. Once you spread users across continents, the same name has to map to different nearby addresses, and the lookup system has to keep working even when a path is partitioned or a data center is down. DNS scales by delegating responsibility in layers, so no one server has to know every name on the internet.
DNS resolution works by having a local component ask a smarter component to do the hard work, then following referrals until it reaches the server that is authoritative for the name. A stub resolver is the small DNS client in the OS that knows how to send a query but does not walk the hierarchy itself, so it forwards to a recursive resolver. The recursive resolver then asks the root, then a TLD server, then the authoritative server, caching results along the way. The end to end exchange looks like this.
In the sequence, the stub sends one query for A www.acme-bakery.example to the recursive resolver, and the recursive resolver does the multi step work. The root server does not return the final address. It returns a referral that points at the .example TLD name servers, typically as NS records plus glue A or AAAA records so the resolver can reach them. The TLD server returns another referral to the authoritative name servers for acme-bakery.example, and only the authoritative server returns the final A or AAAA answer for www.
Design rule
If a resolver cannot reach the next hop, the web request fails before HTTP connects.
Once the resolver reaches the right authoritative server, the answer is expressed as records, each with a type and data format. An authoritative server is a DNS server that is the source of truth for a zone and returns final answers for names in that zone rather than referrals. Most web requests care about A and AAAA, but the same chain and caching rules apply to mail and verification records. The common record types in this scenario are easiest to compare side by side.
The table maps each record type to what the resolver learns. A returns an IPv4 address and AAAA returns an IPv6 address, which are the endpoints the TCP or QUIC connection will target. CNAME returns another name, which forces the resolver to perform additional lookups until it reaches an address record, and that extra chasing adds latency and new failure points. MX names the mail servers for a domain, and TXT carries arbitrary text that is often used for SPF, DKIM, and domain ownership proofs.
DNS performance depends more on caching than on raw server speed, and caching is controlled by TTL on each record. A TTL is the number of seconds a resolver is allowed to reuse a cached answer before it must ask again. There are multiple caches in the path, including the browser, the OS, and the recursive resolver, and each layer can turn a network round trip into a memory lookup when it hits. The tradeoff is that a longer TTL keeps answers around longer even if you changed them. This is where the latency before first byte is decided.
The tuning output shows how changing TTL and hit rate changes the expected number of upstream queries and the added time before the TCP or QUIC handshake can even start. A cache hit at the recursive resolver often saves the whole root to TLD to authoritative walk, while a miss can cost multiple sequential RTTs because the resolver cannot ask the next server until it learns who that server is. With a low TTL like 30 seconds, the resolver refreshes frequently and misses more during traffic spikes. With a high TTL like one day, most users see fast hits but changes such as failover can be delayed until caches expire.
Practical caveat
Short TTL helps agility only if resolvers actually re query instead of retrying stale paths.
DNS problems are often invisible in application logs because they happen before the app server receives anything, and the symptoms vary based on where the chain breaks. A negative caching entry is a cached memory that a name did not exist, typically from an NXDOMAIN response, and it can make a fixed configuration look broken for minutes. Other failures look like timeouts when a resolver cannot reach an authoritative server, or like intermittent errors when delegation points at the wrong servers. The common failure patterns are easiest to spot by where they occur in the chain.
The revealed cards connect user visible behavior to specific layers. An NXDOMAIN tends to be fast because a server confidently says the name does not exist, and negative caching can then keep returning that fast failure until its TTL expires. A timeout tends to be slow because the recursive resolver retries multiple servers and waits for retransmission timers, which stretches the delay before the browser gives up. A lame delegation means the TLD points you to authoritative servers that do not serve the zone, and DNSSEC validation failure happens when signatures do not validate and the resolver rejects answers even though the servers responded. Propagation delays show up when some resolvers still have the old data cached while others have refreshed.
The same web page load scenario becomes predictable when the design choices acknowledge how caching and delegation really behave. Choose recursive resolvers with good anycast coverage and redundant paths, because they sit on the critical path for every connection setup. Set TTLs based on how quickly you need a change to take effect, and avoid long CNAME chains because every extra indirection is another query and another cache that can go stale. The tradeoffs are easiest to evaluate as a small decision matrix rather than a gut feel.
The comparison highlights a stable strategy that many teams converge on, which is splitting TTLs by record role. Keep NS and other delegation records relatively long so the hierarchy stays stable and cacheable, and keep A and AAAA shorter when you need quicker failover between addresses. A redesign is forced when your operational needs change, for example you move from static hosting to frequent traffic steering, or you add DNSSEC and start seeing validation driven outages. DNS looks like a simple lookup, but at global scale it is a distributed cache with explicit staleness controls, and the most defensible designs treat it that way.