Networking Basics

Grasp the network fundamentals that underpin all distributed system architectures.

Last generated

Lesson 3 of 18 available16 practice questions

SPACED REPETITION Β· 16 practice questions

Make this lesson stick.

Try 3 questions now. No account needed. Sample answers aren't saved.

Your server answers in 25 ms. Your user waits 1.8 seconds.

A travel app's home screen calls your API six times: GET /me, then five calls that each need the user id from the first answer (trips, bookings, alerts, points, offers). Each call spends 25 ms on your servers. The servers are in Virginia. Your user is in Sydney.

Put numbers on it. Assume a 200 ms round-trip time (RTT) between Sydney and Virginia, a DNS answer the user's ISP resolver already has cached (20 ms), HTTP/1.1 over TLS 1.3 on a fresh connection, and app code that waits for each call before starting the next. (RTTs in this lesson are assumptions for the exercise. Foundations of System Design has the latency table.)

Step Cost Running total
DNS lookup at the ISP's resolver 20 ms 20 ms
TCP handshake 1 RTT = 200 ms 220 ms
TLS 1.3 handshake 1 RTT = 200 ms 420 ms
Six calls, one after another 6 Γ— (200 + 25) = 1,350 ms 1,770 ms

Your code accounts for 150 ms of that: 8.5%. Almost all the rest is eight round trips across the Pacific.

Predict first: a teammate suggests upgrading the user's connection from 50 Mbps to 500 Mbps. Each response is about 4 KB. How much of the 1,770 ms does that save?

Check your answer

About 3.5 ms. At 50 Mbps a 4 KB response takes 0.64 ms to transmit; at 500 Mbps, 0.06 ms. Six responses save 6 Γ— 0.58 β‰ˆ 3.5 ms. Small requests are bound by round trips, not by bandwidth. A faster link doesn't make light travel faster.

Now remove round trips instead:

Change Time to screen What it removed
Starting point 1,770 ms
Send the five follow-up calls in parallel on one HTTP/2 connection 870 ms 4 round trips
One aggregated endpoint that calls /me, then the other five in parallel, inside the datacenter (0.5 ms RTT there) about 670 ms 1 more round trip
TCP and TLS end at an edge server 15 ms from the user, which keeps a warm connection to Virginia about 315 ms the long trip for both handshakes
Returning user on a connection that is still open about 265 ms the DNS lookup and the handshakes

Nothing on that list made your server faster. Every win came from one habit: count the round trips, then remove them or make them shorter. The last 200 ms is physics. To remove it you move the data (cache it at the edge) or the servers (open a region nearby).

The engineers at Sun who wrote down the "fallacies of distributed computing" in the 1990s put "the network is reliable" and "latency is zero" at the top of the list. This lesson is what you do instead of believing them. Each move starts with a pressure you can hear in the requirements:

# Pressure in the requirements Move What you pay
1 "Fail over to the other region within minutes" Budget the DNS TTL More DNS queries; DNS still can't balance load
2 Users far away; many small calls Count round trips: reuse, multiplex, aggregate, end TLS at the edge Long-lived connections; an aggregation layer to own
3 A late packet is worse than a lost one Pick the transport: TCP, UDP or QUIC With UDP you rebuild what you need of TCP
4 Flaky networks; money moves Make retries safe A key store; retry discipline
5 Lists that change while users scroll Page with a cursor No "jump to page 37"
6 The same bytes go to millions of users Cache at the edge Staleness; purge discipline; leak risk
7 The server must tell clients something now Push with SSE, WebSocket or long polling Stateful connections to scale and drain

This lesson assumes you know what an HTTP request looks like. Load balancers are covered in Core Building Blocks, caches inside your datacenter in Databases & Storage, and pub/sub in Messaging & Queues.

Before any move: know which layer you mean

Interviewers use layer numbers as shorthand. You need three of them:

Name A box working here sees Boxes you will draw
L3, network Source and destination IP addresses Routers; anycast routing
L4, transport IPs and ports; TCP or UDP Network load balancer; firewall rules
L7, application HTTP method, path, headers, cookies Application load balancer; API gateway; CDN edge

One consequence matters more than the names. With HTTPS, a box can see L7 only if it terminates TLS, which means it holds your certificate and decrypts the traffic. An L4 load balancer that passes TLS through untouched cannot route by URL path, cannot cache and cannot add headers. A CDN that caches your pages is decrypting them. Where TLS ends is a design decision, so say it when you draw the diagram.

Why fresh connections are slow even for big files. A new TCP connection doesn't start at full speed. It sends an initial window of about 10 packets (roughly 14 KB) and about doubles that every round trip (slow start). A 1 MB response on a fresh connection needs about seven round trips: 1.4 s at 200 ms RTT, although a 1 Gbps link could move it in 8 ms. A busy warm connection keeps its larger window, so warm connections and nearby edges help large responses too. An idle one may not: by default Linux shrinks the window again after an idle period (the tcp_slow_start_after_idle setting), unless the host turns that off.

Move 1: Budget the DNS TTL

Pressure: "If a region fails, users must reach the other region within five minutes." Or: "We move to the new load balancer on Tuesday."

How a name becomes an address

Browser / OS cache --miss--> Recursive resolver (your ISP's, a public one, or your VPC's)
                                  |  cache miss: start from the top
                                  |--> Root server:          "ask the .com servers"      (referral)
                                  |--> .com TLD server:      "ask ns1.your-dns.example"  (referral)
                                  |--> Your authoritative:   "api.example.com A 203.0.113.10, TTL 60"
                                  v
                    caches the answer for 60 s, hands it back to the OS

The recursive resolver does all the walking. Root and TLD servers only hand out referrals. Resolvers keep those referrals cached for a long time, so a miss usually costs one trip to your authoritative servers. Most queries are single UDP datagrams on port 53. When an answer is too large for UDP, the server marks it truncated and the client asks again over TCP.

The TTL (time to live) on your record tells every cache how long it may reuse the answer. That one number trades two things:

  • Short TTL: caches forget quickly, so a change reaches clients quickly. You pay with more queries to your DNS provider (often billed per query) and more user requests that wait for a lookup.
  • Long TTL: fewer queries and fewer lookup waits, and cached answers keep working for a while if your DNS provider has trouble. You pay with changes that can take the whole TTL to reach everyone.

The failover budget

A resolver that cached the old answer one second before your change keeps it for a full TTL. So for clients that respect TTLs:

worst-case DNS failover = time to detect the failure + TTL

Assume the health check fails the record over after 3 failed checks 10 s apart (30 s), and 50,000 different resolvers ask for your name. Each resolver asks at most once per TTL:

TTL Worst case until every TTL-respecting resolver has the new answer Upper bound on queries to your authoritative servers
30 s 60 s 144 M/day (about 1,700/s)
60 s 90 s 72 M/day (about 830/s)
300 s 330 s 14.4 M/day (about 170/s)
3,600 s about 61 min 1.2 M/day (about 14/s)

A five-minute target rules out 300 s once detection is added. 60 s meets it with room to spare. Going lower buys seconds you may not get: some resolvers and clients keep answers longer than the TTL says.

Three ways DNS failover fails anyway

1. Clients that never ask again. A connection pool resolves the name when it opens a connection, then sends requests over that connection for as long as it lives, whatever the TTL says. AWS's Application Load Balancer, for example, keeps a client's HTTP connection for up to an hour by default (its "HTTP client keepalive duration"), and AWS's docs warn that when traffic is shifted away, clients with open connections may keep sending requests to the old location until they reconnect. Cap connection lifetime on the client, or make the old endpoint close connections politely (Connection: close in HTTP/1.1, a GOAWAY frame in HTTP/2) so clients reconnect and resolve again.

2. Lowering the TTL at the moment of the change.

Predict first: your record has a TTL of 86,400 s (one day). At 10:00 on Tuesday you change the IP and set the TTL to 60 s in the same edit. When is the last TTL-respecting resolver guaranteed to have the new IP?

Check your answer

10:00 on Wednesday. A resolver that cached the record at 09:59 got it with the old 86,400 s TTL and will not ask again for a day. The new TTL only applies to answers handed out after the edit. The recipe for a planned move: lower the TTL at least one old TTL ahead (Monday 10:00 or earlier), switch on Tuesday, keep the old endpoint serving or forwarding until its traffic drains, then raise the TTL again.

3. DNS as the single point of failure. Describing its October 2021 outage, Meta explained how a maintenance command took down its backbone. Its DNS servers, unable to reach the data centers, withdrew their route advertisements as designed, so the internet could no longer reach them and Meta's names stopped resolving, which also broke internal tools the engineers needed for the repair. Your DNS provider is a dependency with its own availability.

⚠️ Negative caching. If something looks up a name before you create it, resolvers cache the "no such name" answer for the negative-caching time in your zone's SOA record. Create records before anyone asks for them.

DNS picks a region; it doesn't balance load

Geo-based, latency-based and weighted DNS answers are good at choosing a region or entry point. They make a poor load balancer:

  • An answer goes to a resolver, not a user, and one resolver can serve thousands of users.
  • Answers are cached for the TTL, so you can't react to load within seconds.
  • DNS can't see how busy a server is, and many clients just take the first address.
  • GeoDNS sees the resolver's location, not the user's. A user on a far-away public resolver can land in the wrong region unless the resolver passes part of the user's address along (EDNS Client Subnet), which some resolvers skip for privacy.

The standard shape is DNS to pick the region, then a load balancer inside it. Anycast is the alternative when minutes are too slow: the same IP address is announced from several locations, internet routing (BGP) delivers each user to a nearby one, and failover means withdrawing a route. No DNS cache has to expire, but you need your own IP ranges and BGP, or a provider that has them.

Say it like this: "DNS with a 60-second TTL picks the region, driven by health checks, and a load balancer inside each region spreads the load. Failover is detection plus TTL, about 90 seconds for well-behaved clients. For clients that hold connections, I cap connection age at a few minutes." Likely follow-up: "What about clients that ignore the TTL?" Cap connection age, retry against a secondary hostname, or put an anycast address in front so the IP never changes.

Your turn: the recovery target drops to 60 seconds. Health checks can run every 5 s, with 3 failures to trip. What TTL do you pick, and is DNS still the right tool?

Check your answer

Detection takes 3 Γ— 5 = 15 s, which leaves at most 45 s for the TTL, so 30 s. That is on paper. Resolvers that stretch short TTLs and pooled connections that never re-resolve make 60 s unreliable. When the target is this tight, stop relying on DNS changes: use an anycast address or a global load balancer whose IP stays the same while the backends behind it change, and have clients retry against a second endpoint. Say what you give up: more cost and a dependency on the provider that owns the anycast network.

Move 2: Count round trips, then cut them

Pressure: users far from your servers, screens that make many small calls, or a service that calls another thousands of times per second.

The price list

Connection Round trips until the first response arrives At 100 ms RTT
TCP + TLS 1.2, new 4 (TCP 1, TLS 2, request 1) 400 ms
TCP + TLS 1.3, new 3 300 ms
TCP + TLS 1.3, resumed with 0-RTT data 2 200 ms
QUIC (HTTP/3), new 2 (transport and TLS handshake together) 200 ms
QUIC, resumed with 0-RTT 1 100 ms
A connection that is already open 1 100 ms

Add a DNS lookup if the name isn't cached. The cheapest handshake is the one you skip.

⚠️ 0-RTT data can be replayed. An attacker who captures it can send it again. So HTTP clients may put only safe requests, such as a GET, in 0-RTT data, and a server that won't risk a replay answers 425 Too Early so the client retries after the handshake (RFC 8470). Note that safe is stricter than idempotent: a PUT or DELETE doesn't qualify.

Four ways to cut round trips

1. Reuse connections. HTTP keep-alive and connection pools pay the handshake once per connection instead of once per request. To size a pool, use Little's law (see Key Concepts & Terminology): connections busy β‰ˆ request rate Γ— time each request holds one. At 3,000 calls/s and 40 ms per call, that is 120 connections on average. Size for the peak and leave headroom.

2. Multiplex. HTTP/1.1 carries one request at a time per connection. Pipelining exists on paper, but browsers don't use it, so they open about six connections per host, and a slow response blocks the requests queued behind it on its connection. That is head-of-line blocking. HTTP/2 interleaves many requests, called streams, on one connection. HTTP/3 does the same over QUIC.

3. Aggregate. When each call needs the previous answer, the client can't parallelize the chain. Move the chain to where round trips are cheap: an aggregation endpoint (a "backend for frontend") or a GraphQL server that makes the calls inside the datacenter at about 0.5 ms each. The cost is one more service to own, and a big composite response that is harder to cache than small ones.

4. Shorten the trip. End TCP and TLS at an edge server near the user: a CDN or your own regional proxy. The handshakes then cost the short RTT, and the edge forwards over a connection to the origin that it keeps warm. Even responses that can't be cached get faster. That was the hook's step from 670 ms to 315 ms.

For calls between your own services, gRPC is the common alternative to JSON over HTTP. It runs on HTTP/2, sends compact binary Protobuf messages, generates typed clients and supports streaming. You give up human-readable payloads, easy browser access and HTTP caching.

Say it like this: "The home screen makes one call to an aggregation endpoint, over a kept-alive HTTP/2 connection that ends at an edge near the user. That turns eight round trips across the Pacific into two short handshakes and one trip to the origin. The cost is an aggregation service we own, and an edge vendor that sees plaintext." Likely follow-up: "The aggregation endpoint calls six services. What if one is slow?" Give each downstream call its own timeout, run them in parallel, and return the screen with the slow section marked as unavailable rather than failing the whole response.

The head-of-line blocking you can't see

Predict first: one HTTP/2 connection carries 30 image downloads in parallel. A single TCP packet belonging to image 7 is lost. Which downloads wait?

Check your answer

Every download whose next bytes come after the lost packet, which is usually all of them. TCP delivers one ordered byte stream. Until the lost packet is retransmitted, at least one more RTT, the kernel won't hand any later bytes to the application, even bytes for other images. HTTP/2 fixed head-of-line blocking at the HTTP layer and inherited it from TCP underneath. HTTP/3 fixes that too: QUIC orders bytes per stream, so only image 7 waits. On a clean network the difference is small. On a lossy mobile link it is large.

What breaks

Ephemeral port exhaustion. Every outgoing TCP connection needs a local port. Linux picks from 32768–60999 by default, which is 28,232 ports (kernel docs). Windows and macOS use the IANA range 49152–65535, which is 16,384. The side that closes a connection first keeps it in the TIME_WAIT state, 60 s on Linux, and the port stays reserved for that destination during the wait. So a proxy that opens a fresh connection per request to one backend IP and port, and closes it itself, can sustain only about 28,232 / 60 β‰ˆ 470 new connections per second before connect() starts failing with EADDRNOTAVAIL. The real fix is pooled keep-alive connections. More backend addresses and kernel settings such as tcp_tw_reuse are the fallback.

The opposite myth: "a server can hold only 65,535 connections, because that's how many ports exist." False. A connection is identified by the protocol plus both endpoints (the 5-tuple: protocol, client IP, client port, server IP, server port). Every client connects to the same port 443, and each connection differs in its client address and port. A server's limits are memory and file descriptors, not ports.

Long-lived connections defeat scale-out. HTTP/2 and gRPC clients open one connection and send everything over it. An L4 load balancer balances connections, not requests. Add five new servers and existing clients keep hammering the old ones. Balance per request at L7, or cap connection age so clients reconnect and spread out.

Your turn: a batch job calls an internal pricing service 2,000 times per second through one load balancer address. For every call it opens a new HTTPS connection and closes it after the response. The RTT is 2 ms and each call takes 10 ms. What breaks first, and what do you change?

Check your answer

Ports. At 2,000 new connections per second, the 28,232 ports are gone in about 14 s. Each closed connection holds its port for 60 s, so from then on connect() fails until ports free up, and the job is capped near 470 new connections per second. Every call also pays two extra round trips of handshakes and the CPU cost of TLS. Use a keep-alive pool instead. By Little's law it needs about 2,000 Γ— 0.010 s = 20 connections on average; size it for the peak.

Move 3: Pick the transport by what hurts more

Pressure: "Voice must feel live." "Game state goes out 30 times a second." "We must never lose a payment."

TCP UDP QUIC
Setup 1 RTT, plus TLS none 1 RTT including TLS 1.3; 0-RTT on resumption
Delivery reliable and in order: one byte stream best effort: may drop, duplicate or reorder reliable and in order within each stream
Message boundaries none; you frame messages each datagram is one message streams; HTTP/3 frames messages
Congestion control yes, in the kernel none; your application must behave yes, in user space
Head-of-line blocking one loss stalls everything behind it none one loss stalls only its own stream
Gets through networks everywhere sometimes blocked by corporate networks runs over UDP, so keep a TCP fallback

The decision comes down to one question: what hurts more, a lost packet or a late one?

  • If losing any byte is wrong (payments, API calls, file transfers), use TCP, or QUIC through HTTP/3.
  • If a packet that arrives late is useless, use UDP. For voice, the ITU's guideline G.114 treats one-way delays up to about 150 ms as fine for most conversations. A retransmission costs at least one more round trip and often arrives after its moment has passed, so the receiver hides the gap instead. For game state, snapshot 41 replaces snapshot 40, so resending 40 is wasted effort.
Workload Choice Why
REST or gRPC calls, payments TCP or QUIC (HTTP/2, HTTP/3) Every byte matters
Video on demand HTTP over TCP or QUIC (HLS, DASH) A buffer of several seconds absorbs lateness
Video and voice calls UDP (RTP, as used by WebRTC) Late audio is useless
Multiplayer game state UDP, adding reliability only for events that matter The newest state wins
DNS lookups UDP, TCP for large answers One small question, one answer
Metrics such as StatsD UDP Losing a sample is fine

⚠️ "Video streaming uses UDP" is a common interview slip. On-demand streaming (Netflix-style or YouTube-style playback) is HTTP over TCP or QUIC. UDP with loss tolerance is for real-time media: calls and conferencing. The YouTube / Netflix Streaming lesson builds on this.

QUIC: UDP carrying a reliable protocol

QUIC runs over UDP but is not unreliable. It rebuilds reliability, per-stream ordering and congestion control in user space, with TLS 1.3 built in. What it buys you:

  • One round trip for transport and encryption together, and 0-RTT on resumption (safe requests such as GET only).
  • Independent streams, so a loss stalls only its own stream.
  • Connection migration. A connection is identified by a connection ID rather than by IP and port, so a phone that moves from Wi-Fi to cellular keeps its connection instead of starting over.

What it costs:

  • Some networks block or throttle UDP, so clients need a TCP fallback. Browsers learn that a server speaks HTTP/3 from an Alt-Svc response header or an HTTPS DNS record, and fall back to HTTP/2 over TCP when UDP fails.
  • It usually costs more CPU per byte than kernel TCP, which has decades of tuning and hardware offload behind it.
  • An L4 load balancer that hashes IPs and ports sends a migrated connection to a different server. To keep migration working, it has to route by connection ID.

Say it like this: "I'd serve the mobile API over HTTP/3 with an HTTP/2 fallback. It saves a round trip on every new connection, a loss on a cellular link stalls only one stream, and connections survive network switches. The price is UDP-aware load balancing and some CPU."

Trap: TCP doesn't know where your messages end

TCP delivers bytes, not messages. Two writes can arrive as one read, and one write can arrive as two:

import socket, struct, time

server = socket.create_server(("127.0.0.1", 0))      # any free port
client = socket.create_connection(server.getsockname())
conn, _ = server.accept()

client.sendall(b"hello")
client.sendall(b"world")
time.sleep(0.05)                                     # let both writes arrive
print(conn.recv(1024))                               # b'helloworld': two messages, one read

def send_msg(sock, payload: bytes):
    sock.sendall(struct.pack("!I", len(payload)) + payload)   # 4-byte length, then body

def recv_exactly(sock, n):
    buf = b""
    while len(buf) < n:
        chunk = sock.recv(n - len(buf))
        if not chunk:
            raise ConnectionError("peer closed mid-message")
        buf += chunk
    return buf

def recv_msg(sock):
    (length,) = struct.unpack("!I", recv_exactly(sock, 4))
    return recv_exactly(sock, length)

send_msg(client, b"hello")
send_msg(client, b"world")
print(recv_msg(conn), recv_msg(conn))                # b'hello' b'world'

HTTP marks where a body ends with Content-Length or chunked encoding. WebSocket and gRPC frame messages for you. You only write framing yourself if you design your own protocol on raw TCP, and then you must. UDP keeps message boundaries but has its own limit: keep datagrams small (about 1,200 bytes), because a datagram that gets split into IP fragments is lost if any fragment is lost.

Your turn: a delivery app shows each courier's position on the customer's map, one update per second, and bills restaurants by the distance each courier drove. Couriers are on mobile networks. Would you send positions over UDP?

Check your answer

Probably not, and the reason is the numbers, not a slogan. At one update per second, a TCP retransmission that delays an update by one round trip is invisible on a map, so a WebSocket or a plain HTTPS stream is simpler and gets through every network. UDP pays off when updates are frequent and the latency budget is tight, as with voice or game state at 20 to 60 packets per second. Billing is a different requirement: every point matters, so the courier app uploads the recorded track reliably over HTTPS, in batches, with an idempotency key (next move) so a retried upload isn't counted twice.

Move 4: Make retries safe

Pressure: mobile clients on flaky networks, timeouts between services, anything that moves money or creates records.

First, the shape of a REST API. Resources are nouns, and HTTP methods are the verbs:

POST   /orders                      create an order                           -> 201 Created, Location: /orders/o_81
GET    /orders/o_81                 read it                                   -> 200
GET    /orders?cursor=...           list my orders, paged                     -> 200 with next_cursor (Move 5)
PATCH  /orders/o_81                 change delivery notes                     -> 200
POST   /orders/o_81/cancellation    cancel it (an action, as a new resource)  -> 201
DELETE /users/me/addresses/a_7      remove a saved address                    -> 204

An action that doesn't fit the verbs becomes a sub-resource, like cancellation or refunds, rather than an endpoint named /cancelOrderAndRefund. Keep DELETE for things that really go away: a cancelled order still exists, with its status and history.

The question a timeout can't answer

App                       Payments API              Database        Payment provider
 |--- POST /payments ---------->|                       |                   |
 |                              |--- insert payment --->|                   |
 |                              |--- charge card -------------------------->|
 |                              |<-- ok ------------------------------------|
 |    x<---- 201 Created -------|  (response lost in a tunnel)              |
 | (times out after 3 s)        |                       |                   |
 |--- POST /payments (retry) -->|--- insert AGAIN ----->|                   |
 |                              |--- charge card AGAIN -------------------->|   two charges

Predict first: the app's POST timed out after 3 seconds. Did the payment happen?

Check your answer

You can't tell. A timeout, a dropped connection or a 504 Gateway Timeout means "no answer", not "not done". The request may have died before reaching the server, failed while it ran, or succeeded with the answer lost on the way back. So a retry is safe only if doing the operation twice has the same effect as doing it once. That property is called idempotent.

Which methods can be retried

Method Typical use Safe (read-only)? Idempotent?
GET, HEAD read yes yes
PUT replace the resource at a known URL no yes
DELETE remove no yes
POST create something, or run an action no no
PATCH partial update no not guaranteed

Safe and idempotent are defined in RFC 9110; PATCH itself comes from RFC 5789. Idempotence is about the effect on the server, not the response: a second DELETE may return 404 instead of 204, but the resource is gone either way. The definitions are also promises your server has to keep. A GET /jobs/run that starts a job breaks the promise, and caches, link prefetchers and retrying clients will start it for you.

Two ways to make "create" idempotent

  1. Let the client name the resource. The client generates an id (a UUID) and sends PUT /orders/7c1e…. A retry targets the same URL, and a second PUT of the same content changes nothing.
  2. Send an idempotency key with the POST. The client generates a random key once per user action, not once per attempt, and sends it as Idempotency-Key: …. The server saves the key together with its response and replays that response for repeats. Stripe's API works this way (its docs say keys may be pruned once they are 24 hours old), and an IETF draft describes the header.

Here is the server side, with SQLite standing in for your database:

import json, sqlite3, uuid

db = sqlite3.connect(":memory:", isolation_level=None)   # we issue BEGIN/COMMIT ourselves
db.executescript("""
CREATE TABLE orders (id TEXT PRIMARY KEY, user_id TEXT, amount_cents INTEGER);
CREATE TABLE idempotency_keys (
    user_id TEXT, key TEXT, request TEXT, status INTEGER, body TEXT,
    PRIMARY KEY (user_id, key));
""")

def create_order(user_id, key, payload):
    request = json.dumps(payload, sort_keys=True)          # a real service stores a hash
    db.execute("BEGIN IMMEDIATE")                          # one writer at a time
    try:
        saved = db.execute(
            "SELECT request, status, body FROM idempotency_keys WHERE user_id = ? AND key = ?",
            (user_id, key)).fetchone()
        if saved:
            db.execute("ROLLBACK")
            if saved[0] != request:
                return 422, {"error": "key reused with a different request"}
            return saved[1], json.loads(saved[2])          # replay the first answer
        order_id = str(uuid.uuid4())
        db.execute("INSERT INTO orders VALUES (?, ?, ?)",
                   (order_id, user_id, payload["amount_cents"]))
        body = {"order_id": order_id}
        db.execute("INSERT INTO idempotency_keys VALUES (?, ?, ?, ?, ?)",
                   (user_id, key, request, 201, json.dumps(body)))
        db.execute("COMMIT")                               # order and key: all or nothing
        return 201, body
    except Exception:
        db.execute("ROLLBACK")
        raise

key = str(uuid.uuid4())                  # made once, before the first attempt
first = create_order("u1", key, {"amount_cents": 4999})
retry = create_order("u1", key, {"amount_cents": 4999})   # the first response was lost
print(first == retry)                                      # True: same order id
print(db.execute("SELECT COUNT(*) FROM orders").fetchone()[0])   # 1
print(create_order("u1", key, {"amount_cents": 1})[0])     # 422

Four details decide whether this works in production:

  • Atomic. The key row and the order row commit in one transaction. Save the key in a separate step after the order, and a crash between the two leaves an order with no key. The retry then creates a second order.
  • Scoped to the caller, (user_id, key), so one client can't replay another's response.
  • Checked against the request. The same key with a different body is a client bug. The IETF draft suggests 422; Stripe rejects it too.
  • Work that calls other services. You can't hold a database transaction open around a slow call to a payment provider. Record the key as "in progress" first, answer duplicates that arrive meanwhile with 409 Conflict (the draft's suggestion), and pass the same key downstream so the provider deduplicates its own side. That 409 is temporary: the client backs off and retries with the same key, and once the first request finishes it gets the saved response.

Status codes: what the client does next

Code Meaning Client's next move
200, 201, 204 done (201 carries a Location) carry on
202 Accepted queued, not done yet poll a status URL, or wait for a push
304 Not Modified your cached copy is still good use it (Move 6)
301, 308 / 302, 307 moved permanently / temporarily; 307 and 308 keep the method follow it; caches may store permanent redirects
400, 422 the request is invalid fix it; a retry won't help
401 / 403 not signed in / not allowed sign in / stop
404 no such resource stop
409 Conflict, "key still in progress" a keyed request with the same key is still running back off, retry with the same key
409 Conflict, any other clashes with the current state re-read, then decide
412 Precondition Failed your If-Match version is stale re-read, merge, try again
429 Too Many Requests rate limited wait as Retry-After says, then retry
500 unexpected server error, often a bug usually not retried; if at all, once and only if safe
502, 503, 504 bad gateway / unavailable / gateway timeout retry with backoff if safe; after a 504 the outcome is unknown

Rate limiting and 429 get their own treatment in Design URL Shortener & Rate Limiter. The 301-versus-302 choice matters most in the URL Shortener design.

The rule: retry only when the failure might be temporary and a repeat is harmless. Temporary means timeouts, dropped connections, 429, 502, 503 and 504, plus a 409 that says your idempotency key is still in progress. Harmless means an idempotent method or an idempotency key. An API that uses 409 for both meanings must tell them apart, for example with an error code in the body, so clients know which one to retry.

Retries without a stampede

import random, time

RETRYABLE = {429, 502, 503, 504}

def call_with_retries(send, attempts=4, base=0.1, cap=5.0, sleep=time.sleep, keyed=False):
    """send() returns (status, retry_after_seconds or None), or raises on timeout.
    Only for idempotent requests, or requests that carry an idempotency key.
    keyed=True: the request carries a key, and this API uses 409 only for "this key is
    still in progress" (real conflicts come back as 412 or 422), so a 409 is temporary."""
    retryable = (RETRYABLE | {409}) if keyed else RETRYABLE
    for attempt in range(attempts):
        try:
            status, retry_after = send()
        except (TimeoutError, ConnectionError):
            status, retry_after = None, None       # outcome unknown
        if status is not None and status not in retryable:
            return status                           # success, or an error a retry won't fix
        if attempt == attempts - 1:
            raise RuntimeError(f"gave up after {attempts} attempts")
        delay = random.uniform(0, min(cap, base * 2 ** attempt))   # full jitter
        if retry_after is not None:
            delay = max(delay, retry_after)         # the server knows better
        sleep(delay)

# Demo, a keyed POST: a timeout, then 409 (the first attempt is still running),
# then a 503 asking for 1 s, then the saved 201.
replies = iter([TimeoutError(), (409, None), (503, 1.0), (201, None)])
def send():
    r = next(replies)
    if isinstance(r, Exception):
        raise r
    return r

random.seed(7)
waits = []
print(call_with_retries(send, sleep=waits.append, keyed=True))   # 201
print([round(w, 3) for w in waits])                              # [0.032, 0.03, 1.0]
  • Exponential backoff doubles the ceiling on each wait: 0.1, 0.2, 0.4, 0.8 s and so on, capped at 5 s. Full jitter picks a random wait between zero and that ceiling, so thousands of clients that failed together don't retry together. In AWS's simulations, full jitter did much less work than backoff without jitter.
  • Put a timeout on every remote call, shorter than your caller's timeout. Without one, a slow dependency ties up all your threads or connections.
  • Retry at one layer. If the app, the gateway and the order service each make 3 attempts, one tap can become 3 Γ— 3 Γ— 3 = 27 calls to the inventory service at the bottom, arriving just when it can least handle them. Google's SRE book recommends retrying only in the layer directly above the failing one, and capping retries at about 10% of a client's requests (a retry budget).
  • Circuit breaker. After repeated failures, stop calling for a while and fail fast, then let a few test requests through.

Plot twist: retry-safe is not race-safe

Alice and Bob both open document version 7. Alice saves with a PUT, then Bob saves with a PUT. Both requests are idempotent, and Bob's silently wipes out Alice's change. This is a lost update, and retry safety doesn't prevent it. Versions do:

Step Alice Bob Server version
1 GET returns ETag: "v7" GET returns ETag: "v7" v7
2 PUT with If-Match: "v7" returns 200, ETag: "v8" v8
3 PUT with If-Match: "v7" returns 412 Precondition Failed v8
4 re-reads v8, merges, PUT with If-Match: "v8" returns 200 v9

Say it like this: "POST /payments requires an Idempotency-Key. The key and the payment commit in one transaction, and a retry gets the saved response. Clients retry timeouts, 429, 502, 503 and 504, and a 409 for a key still in progress, with capped exponential backoff and full jitter, and only one layer retries, so retries don't multiply." Likely follow-up: "What if the key store is down?" For money, fail closed: reject the write rather than risk charging twice.

Your turn: the operation changes to "set my notification preferences". The app sends PUT /users/me/preferences with the complete preferences object. Does this need an idempotency key? What can still go wrong?

Check your answer

No key needed. A PUT that replaces the whole object is idempotent: sending it twice leaves the same state, so the app can retry freely. What can still go wrong is concurrency. The user changes one setting on a phone and another on a laptop. Each device PUTs its full object, and the second silently undoes the first. Fix it with If-Match and a 412 on conflicts, or send only the changed fields, for example a PATCH that sets one field to a value, which is also safe to repeat.

Move 5: Page with a cursor

Pressure: "infinite scroll over a feed that gets new posts every second", "export 50 million rows through the API".

Offset pagination looks innocent: GET /posts?offset=40&limit=20 becomes ORDER BY created_at DESC LIMIT 20 OFFSET 40. It has two problems.

  1. Cost grows with depth. The database has to walk past every row before the offset. The first page reads 20 rows. The page at offset 1,000,000 walks past a million rows to return 20.
  2. The list moves under you. Watch what one new post does:
# Posts sorted newest first. Each post: (created_at, id).
feed = [(100 + i, i) for i in range(10, 0, -1)]      # ids 10..1

def by_offset(items, offset, limit):
    return items[offset:offset + limit]

def by_cursor(items, cursor, limit):
    if cursor is not None:
        items = [p for p in items if p < cursor]      # strictly after the last one seen
    page = items[:limit]
    return page, (page[-1] if page else None)

ids = lambda page: [p[1] for p in page]

page1 = by_offset(feed, 0, 3)
c_page1, cursor = by_cursor(feed, None, 3)

feed.insert(0, (111, 11))                             # a new post arrives

print("offset page 1:", ids(page1), "page 2:", ids(by_offset(feed, 3, 3)))
print("cursor page 1:", ids(c_page1), "page 2:", ids(by_cursor(feed, cursor, 3)[0]))
offset page 1: [10, 9, 8] page 2: [8, 7, 6]
cursor page 1: [10, 9, 8] page 2: [7, 6, 5]

The new post pushed everything down one slot, so offset 3 now points at post 8 again: a duplicate. Deletions do the opposite and make posts vanish between pages.

Cursor pagination (also called keyset pagination) returns an opaque next_cursor that encodes the sort key of the last item, and the next query starts strictly after it:

SELECT id, created_at, body
FROM posts
WHERE (created_at, id) < (:last_created_at, :last_id)   -- PostgreSQL row comparison
ORDER BY created_at DESC, id DESC
LIMIT 21;                                               -- one extra row tells you has_more
{ "items": ["...20 posts..."], "next_cursor": "eyJ0IjoxNzI3MTAwMDAwLCJpZCI6OTA0Mn0" }

With an index on (created_at, id), every page costs one index seek plus 21 rows, at any depth, and new posts at the top can't shift what comes after the cursor. The cursor above is just {"t":1727100000,"id":9042} in base64. Keep it opaque so you can change the encoding later, and never let clients build cursors themselves.

What you give up: there is no "jump to page 37" and no cheap total count, and a cursor works for one sort order only. Infinite scroll, sync and exports don't need any of those.

Your turn: someone "simplifies" the cursor to created_at alone, and timestamps have one-second resolution. Posts 8, 7 and 6 share the same second, and a page ends at post 7. What does the next page show with WHERE created_at < :cursor? And with <=?

Check your answer

With <, post 6 vanishes: it has the same timestamp as the cursor, so it is not strictly older. With <=, posts 8 and 7 come back as duplicates. A cursor has to identify a position uniquely, so add a unique tiebreaker, (created_at, id), and use it in both the ORDER BY and the comparison.

Say it like this: "The feed API uses cursor pagination on (created_at, id) with an opaque cursor. Every page costs the same at any depth and stays stable while new posts arrive. We give up random page access, which infinite scroll doesn't need." Likely follow-up: "What about a ranked feed, where order isn't a column?" Save the ranked list of ids for the session and page through that. The Twitter/Instagram News Feed lesson goes deeper.

Move 6: Cache at the edge

Pressure: the same bytes go to many users (images, scripts, video segments, public pages), users are far away, or the origin can't take the load.

A CDN is a fleet of caching reverse proxies close to users. On a hit, the edge answers by itself: a short round trip and no work for your origin. On a miss, it fetches the object from your origin, stores it under a cache key (by default, roughly the URL), and answers.

Users reach a nearby edge in one of two ways. With anycast, the same IP address is announced from every location and internet routing picks a close one (Cloudflare works this way). With DNS-based steering, the CDN's DNS hands out a different edge address depending on where the query comes from. Either way, "nearest" is approximate.

Pull or push

Pull Push
How content gets to the edge The edge fetches from your origin on the first miss You upload it to the CDN's storage ahead of time
Good for Most sites: many objects, popularity unknown Big files known in advance: game patches, a launch-day video
Cost The first request at each edge is slow; a new popular object can stampede the origin You manage uploads, storage and deletion

The headers that decide everything

Your origin controls caching with response headers (HTTP caching is specified in RFC 9111):

Header Meaning
Cache-Control: max-age=N any cache may reuse the response for N seconds
s-maxage=N the lifetime for shared caches such as a CDN; overrides max-age there
public / private public: shared caches may store it, even for a request that carried Authorization. private: only the user's own browser may store it, never a CDN
no-cache may be stored, but must be revalidated with the origin before every reuse
no-store must not be stored anywhere
immutable the content at this URL will never change, so don't revalidate
stale-while-revalidate=N serve a stale copy for up to N s while fetching a fresh one in the background
stale-if-error=N serve a stale copy for up to N s if the origin is failing
ETag with If-None-Match revalidation: the origin answers 304 Not Modified with no body if nothing changed
Vary: Accept-Encoding the named request headers become part of the cache key

⚠️ no-cache doesn't mean "don't cache". It means "check before reusing". The directive that forbids storage is no-store. Support for the stale-* directives varies by CDN.

A policy for each kind of resource:

Resource Header Why
app.3f9a1c.js, logo.7b2e.png (content hash in the name) public, max-age=31536000, immutable New content gets a new URL, so the old URL never goes stale
Product page HTML, the same for everyone public, max-age=0, s-maxage=60, stale-while-revalidate=30 Changes show within about 90 s; the origin sees one refresh per minute per edge, not every view
Home page HTML that must change right after a deploy no-cache with an ETag Every view checks with the origin, but most checks are cheap 304s with no body
/api/cart, account pages private, no-store Personalized, so never in a shared cache

Compress text responses (JSON, HTML, JavaScript) with gzip or Brotli. Text usually shrinks a lot. JPEG, MP4 and zip-based formats are already compressed and barely shrink at all.

Getting changes out: version URLs, purge for emergencies

  • Versioned URLs are the default. Put a content hash in the file name, cache it for a year, and change the reference to it. The short-lived HTML points to the new name. There is nothing to purge, and the old and new versions coexist during a deploy, so no page gets new HTML with old JavaScript.
  • Purge (invalidation) tells the edges to drop a URL. Use it for mistakes and for URLs you can't version, such as the HTML itself or /robots.txt. It takes seconds to minutes to reach every edge depending on the CDN, some providers charge per path, and purging a hot object sends a burst of misses to your origin.

The offload estimate: think in miss ratio

Origin traffic = edge traffic Γ— (1 βˆ’ hit ratio). Take a marketplace that serves 80,000 image requests per second at peak, averaging 120 KB. That is 80,000 Γ— 120,000 bytes Γ— 8 = 76.8 Gbps leaving the edges.

Hit ratio Origin egress Origin requests
90% 7.68 Gbps 8,000/s
95% 3.84 Gbps 4,000/s
99% 0.77 Gbps 800/s

Predict first: going from a 95% to a 99% hit ratio sounds like a 4% improvement. How much less traffic reaches the origin?

Check your answer

Five times less. The origin sees the misses, and the miss ratio falls from 5% to 1%. That is why teams chase the last few points: normalizing cache keys, adding a shield tier, and giving content-hashed files very long lifetimes.

What breaks

  • Stampede on a miss. A hot object expires or is purged, and 5,000 requests reach an edge in the same instant. If each one misses through to the origin, the origin gets 5,000 identical requests. Fixes: request collapsing, where the edge sends one request to the origin and holds the others until it returns (many CDNs support this), an origin shield, a middle cache tier so only one location asks the origin instead of every edge, and stale-while-revalidate.
  • A personalized page in a shared cache. A rule like "cache everything for 5 minutes" meets a page that depends on the session cookie, and the first user's account page is served to everyone at that edge. The origin must mark such responses private or no-store, the CDN must honor that, and the cache key must include everything the response depends on.
  • Fragmented cache keys. Tracking parameters (?utm_source=…), Vary: User-Agent or Vary: Cookie split one object into thousands of keys, and the hit ratio collapses. Drop irrelevant query parameters from the key, and vary only on headers that really change the body.
  • Origin outage. stale-if-error keeps serving the last good copy while the origin is down. For a news site, a slightly stale article beats an error page.

"Uncacheable" doesn't mean "no CDN"

Even a fully personalized API gains from an edge: handshakes end 15 ms from the user instead of 200 (the hook's step from 670 to 315 ms), the edge reuses warm connections to your origin, it can speak HTTP/3 to phones, and it absorbs attack traffic. The cost: another vendor in the path, one that sees your decrypted traffic.

Say it like this: "Static assets get content-hashed URLs with a one-year immutable lifetime, so deploys never need a purge. Shared HTML gets s-maxage=60 with stale-while-revalidate. Anything personalized is private, no-store. At a 99% hit ratio the origin sees 1% of the traffic, and an origin shield plus request collapsing protect it from stampedes."

Your turn: product prices on the marketplace change often and must be correct within 10 seconds, and product pages get 20,000 views per second. What do you cache, and for how long?

Check your answer

Split the page by how fast each part changes. The page shell (layout, description, photos) is shared and changes rarely, so give it a long s-maxage and purge it by tag when a seller edits the product (many CDNs can purge every URL tagged with a product id). The price comes from a small endpoint with max-age=0, s-maxage=5: with request collapsing, the origin sees at most one request per product every 5 seconds per edge or shield location, not 20,000 per second, and a price is never more than about 5 s stale plus the fetch time. The other option, caching the whole page for a few seconds (safely under 10), also works, but it re-fetches the heavy page every few seconds just to refresh one number.

Move 7: Push without polling yourself to death

Pressure: "notify the user when…", "live scores", "show the other person's cursor".

HTTP is request and response: the server can't speak first. There are four ways around that:

Short polling Long polling Server-Sent Events (SSE) WebSocket
How it works The client asks every N seconds The client asks; the server holds the request until there is news or a timeout; the client asks again One long HTTP response that streams events (text/event-stream) An HTTP upgrade (extended CONNECT on HTTP/2) to a two-way connection that stays open
Direction client pulls client pulls, but it feels like push server to client both ways
Delay up to N seconds immediate while a request is waiting; a burst waits a round trip immediate immediate
After a drop next poll next poll, which must carry a cursor built in: the browser reconnects and sends Last-Event-ID you build it
Through proxies easy easy (watch timeouts) plain HTTP proxies must allow the upgrade
Idle cost per client a request every N s a request per hold period one open connection one open connection

Here is what each costs with 1,000,000 clients connected:

Approach Load while nothing happens
Short polling every 5 s 200,000 requests/s, almost all answered "nothing new"
Long polling, 30 s hold about 33,000 requests/s, plus one per delivered event
SSE or WebSocket no requests; 1,000,000 open connections (20 GB at an assumed 20 KB each, including kernel buffers) and, with a heartbeat every 30 s, about 33,000 tiny frames/s

Holding that many mostly idle connections is a job for event-driven servers, where one thread watches thousands of sockets through the kernel's readiness APIs (epoll on Linux, kqueue on macOS and BSD), as Nginx and Node.js do. A thread per connection runs out of memory and drowns in context switches long before a million.

Which one to choose:

  • SSE when data flows from server to client: notifications, live scores, dashboards, progress bars, streamed AI answers. It is plain HTTP, so it works with cookie-based auth, your proxies and HTTP/2, and the browser reconnects by itself. Over HTTP/1.1, browsers allow only six connections per domain across all tabs, so the seventh tab of your app starves. Serve SSE over HTTP/2. The browser's EventSource can't send an Authorization header, so an API that uses bearer tokens needs a cookie or a fetch-based stream reader instead. The client still sends data with ordinary POSTs.
  • WebSocket when both sides talk often and quickly: chat, collaborative editing, multiplayer games in the browser. You build reconnection, resume, heartbeats and acknowledgements yourself. The Chat Application (WhatsApp) lesson does exactly that.
  • Long polling as the fallback when a proxy breaks streaming, or when events are rare.
  • Webhooks (you POST to the customer's URL) for server-to-server notifications. Phones whose app is in the background need the platform's push service (APNs, FCM), because your connection is gone.

⚠️ HTTP/2 "server push" is not a way to push events. It pushed files into the browser's cache, and browsers have dropped it (Chrome in version 106, Firefox in version 132).

What breaks when connections live for hours

  • Idle timeouts. Load balancers and NAT gateways close quiet connections. An AWS Application Load Balancer, for example, closes a connection after 60 s of silence by default (its "connection idle timeout", a different setting from the one-hour keepalive in Move 1). Send a heartbeat (a WebSocket ping, or an SSE comment line) more often than the shortest timeout on the path, for example every 30 s.
  • Deploys and crashes cause reconnect storms. A server holding 100,000 connections restarts, and 100,000 clients reconnect in the same second, each with a TLS handshake and a "what did I miss?" query. Spread them out: clients wait a random delay before reconnecting (jitter, growing exponentially if they fail again), the server drains by closing connections gradually over minutes before it stops, and the catch-up is cheap because it reads from a last event id instead of running a heavy query. Spread over 30 s, that is about 3,300 reconnects per second instead of 100,000 at once.
  • Finding the right server. To push to user 42 you have to know which of your ten servers holds user 42's connection. Keep a registry from user to server, or use pub/sub, where each server subscribes to the channels of its connected users (Messaging & Queues).
  • The 65,535 myth from Move 2 doesn't apply here: one server can hold far more connections than that on port 443.

Say it like this: "Scores only flow down, so I'd use SSE over HTTP/2 with event ids. Clients resume with Last-Event-ID after a drop and reconnect with jittered backoff, so a deploy doesn't stampede us. Each server subscribes to the matches its clients watch. What we give up is a two-way channel, so viewers' rare actions go over normal POSTs."

Your turn: the live-score page now lets viewers post reactions, about one per viewer per minute, with 1,000,000 viewers. Do you switch to WebSockets?

Check your answer

No. One reaction per viewer per minute is 1,000,000 / 60 β‰ˆ 16,700 POSTs per second, which an ordinary HTTP API tier handles, while SSE keeps carrying the scores down. WebSockets earn their extra cost (your own reconnect, resume and acknowledgement logic, and proxy quirks) when the client sends often and every message is latency-sensitive, like cursor positions 20 times a second in a collaborative editor.

Final round: no label on the problem

Real prompts don't say which move they need. Read the requirements, find the pressure, pick the move, and name what it costs.

Challenge 1: The ticket drop

A concert sells 40,000 tickets at 10:00. In the first minute, 300,000 fans tap Buy on their phones, many on crowded mobile networks. Nobody may be charged twice. A fan whose app loses signal must still find out whether they got a ticket. The payment provider accepts at most 2,000 requests per second.

  1. What does the Buy request look like, and what makes a retry safe?
  2. What should the app do when the request times out?
  3. How does the fan learn the result?
Hint

300,000 taps in 60 seconds is 5,000 per second, which is more than the payment provider accepts.

Check your answer
  • Request: POST /orders with an Idempotency-Key that the app generates when the fan taps Buy and reuses on every retry. The server stores the key with the order in one transaction and passes it on to the payment provider.
  • Load: 5,000 per second arrive, but the provider takes 2,000, so the API answers 202 Accepted with a status URL and puts the order in a queue that feeds payments at the provider's rate (Messaging & Queues). Charging the 40,000 winners takes about 20 s at that rate; everyone after them gets "sold out" without touching payments. When even the queue must shed load, the edge returns 429 with Retry-After.
  • Timeouts: the app retries the same POST with the same key, using capped exponential backoff with full jitter, and treats a 409 "still in progress" as a reason to wait and retry. It never generates a new key for the same tap.
  • Result: the app polls the status URL, or listens on SSE while it is open. After losing signal, it asks for the order by its key. A duplicate POST gets the saved answer, so it can never create a second order.
  • What you give up: fans wait in a queue instead of getting an instant answer, and you run a key store and a queue.

Challenge 2: The status page that must survive the outage

Your SaaS company's status page must stay reachable while the main platform is down. Traffic is normally 50 requests per second, but during an incident it jumps to 5,000 requests per second, all reading the same page.

Check your answer
  • Hosting: a static page on object storage behind a CDN, on infrastructure that shares nothing with the platform: a separate domain, and ideally a separate DNS provider. If your main DNS provider fails, the status page must still resolve.
  • Caching: everyone gets the same page, so public, max-age=30, stale-while-revalidate=30, stale-if-error=86400. The hit ratio is close to 100%. The origin sees about one refresh per 30 s per edge location, not 5,000 per second, and an origin failure serves the last good copy instead of an error.
  • Updates: publish by writing a new object. Readers see it within about a minute, which is fine for a status page.
  • What you give up: a second stack to own and pay for, and up to a minute of staleness.

Challenge 3: The live auction

A hot auction item has 20,000 people watching it. Every watcher must see a new bid within 500 ms. Bidders bid from the same page, a few times per minute at most. A bid whose request times out must not be placed twice, and a bid made against an out-of-date price must be rejected, not silently accepted.

Check your answer
  • Down: SSE carries price updates to watchers, with event ids so a reconnecting watcher can catch up. A WebSocket is also acceptable, but bids are rare, so it isn't needed.
  • Up: POST /items/42/bids with an Idempotency-Key. The bid carries the price the bidder saw, and the server accepts it only if that is still the current price, like an If-Match precondition. Otherwise it returns 412 Precondition Failed with the new price, which the app shows instead of retrying.
  • Fan-out: the service that accepts bids publishes each new price to a channel. Every connection server subscribed to that item forwards the price to its watchers. 20,000 watchers spread over a few servers is a small fan-out.
  • Failure: heartbeats keep idle connections alive through load balancer timeouts. After a deploy, jittered reconnects plus Last-Event-ID resume mean nobody misses a price.

Challenge 4: Photos on a lossy network

Your photo app's fastest-growing market is 250 ms RTT from your nearest region, and mobile networks there lose about 2% of packets. The gallery screen loads a list of 30 photos, then their thumbnails.

Check your answer
  • Thumbnails: served from a CDN with content-hashed URLs and max-age=31536000, immutable, so a returning user loads them from the browser cache or a nearby edge.
  • Protocol: HTTP/3 at the edge, with an HTTP/2 fallback. A new connection costs one round trip less, a lost packet stalls one thumbnail instead of all 30, and connections survive switches between Wi-Fi and cellular.
  • List API: one call returns the page of 30 items with their thumbnail URLs, using cursor pagination, marked private. The edge can't cache it, but it still goes through the edge so the handshakes stay local. Prefetch the next page.
  • Later: if the market keeps growing, a region nearby removes the last 250 ms. Name the cost: data residency and replication questions (Databases & Storage).

Cheat sheet: pressure β†’ move β†’ cost

When the requirements say… Reach for And say what you give up
"Fail over within N minutes" TTL ≀ N minus detection time; capped connection age; anycast when N is tiny More DNS queries; DNS doesn't balance load
Users far away, many small calls Reuse connections, HTTP/2 or HTTP/3, aggregate chains, end TLS at the edge An aggregation layer; the edge sees plaintext
A late packet is useless UDP, adding only the reliability you need You handle loss, ordering and congestion yourself
Lossy mobile networks HTTP/3 over QUIC with a TCP fallback UDP-aware load balancing; CPU
Clients retry; money moves Idempotency keys committed with the write; backoff with full jitter at one layer A key store; fail closed if it's down
Concurrent edits ETag plus If-Match, 412 on conflict Clients must re-read and merge
A list that changes while users page A cursor on (sort key, unique id) No random page access, no cheap total
The same bytes to many users CDN, hashed URLs marked immutable, s-maxage for shared HTML Staleness; purge discipline; leak risk
A hot object expires Request collapsing, an origin shield, stale-while-revalidate Brief staleness
Server-to-client updates SSE over HTTP/2; WebSocket if both directions are busy; long polling as a fallback Stateful connections: heartbeats, draining, routing

Before moving on, answer this aloud, as you would in an interview, in about three minutes: "A user in Sydney opens our app for the first time. Walk me through what the network does." Name each step (DNS, TCP, TLS, the requests, the edge), how many round trips each costs, and the move that cuts it. Then pick one move and state what it gives up. Finish with the follow-up an interviewer will ask next: "And what happens when that server dies?" If you can answer both without notes, you have the networking half of every design that follows.

Next: Core Building Blocks for L4 and L7 load balancers and the stateless service tier, Databases & Storage for caching inside your datacenter, and Messaging & Queues for the pub/sub behind real-time fan-out.