Requirements Gathering
Learn to ask the right clarifying questions to scope functional and non-functional requirements.
SPACED REPETITION Β· 16 practice questions
Make this lesson stick.
Try 3 questions now. No account needed. Sample answers aren't saved.
or sign in to practice all 16One sentence you didn't ask for
"Design a leaderboard for our mobile game." You have seen leaderboards before. Within a minute you have a table, an index and a query:
CREATE TABLE scores (player_id BIGINT PRIMARY KEY, score INT NOT NULL);
CREATE INDEX scores_by_score ON scores (score DESC);
SELECT player_id, score FROM scores ORDER BY score DESC LIMIT 100;
That is a good design. The index keeps scores in order, so the top-100 query reads about 100 index entries whether the table holds a thousand players or a hundred million.
Then the interviewer adds one sentence: "Every player sees their own global rank on the home screen, updated after each match. We have 100 million players."
Predict first: which part of your design breaks, and why?
Check your answer
The rank lookup. A player's rank is one plus the number of players with a higher score:
SELECT COUNT(*) + 1 FROM scores WHERE score > :my_score;
The index keeps entries in order, but a PostgreSQL or MySQL B-tree does not store how many entries sit under each branch. So the database walks every entry above you and counts it. A player in the middle of 100 million players costs about 50 million steps. The top-100 query was bounded. This one grows with the player's rank.
Put numbers on it. Assume 20 million players play on a given day, each opens the home screen 5 times, and the evening peak is 3Γ the daily average:
| Quantity | Value |
|---|---|
| Rank lookups | 100M a day: 1,157/s on average, 3,472/s at peak |
| Index entries counted per mid-table lookup | about 50 million |
| Index entries counted per second at peak | about 170 billion |
In a quick in-memory SQLite test on a laptop, counting one million index entries took about 16 ms on one core. At that speed, one mid-table lookup takes most of a second, and the peak would need thousands of cores. The design was right for the prompt you heard and wrong for the one the interviewer meant. The only difference was one sentence, and a question would have brought it out.
Better still, one more question splits the design three ways: "Does the player need the exact rank, or is 'top 3%' enough? And how fresh must it be?"
| Answer | Design | What you give up |
|---|---|---|
| Only the top 100 matter | The SQL table above | Nothing: it is already right |
| Exact rank, live | An in-memory sorted structure. A Redis sorted set answers ZREVRANK in O(log N) time, and logβ of 100 million is about 27 |
All players held in memory, and a second copy of every score to keep in sync |
| "Top 3%", hourly | A batch job computes score percentiles every hour; a lookup is a binary search over 100 cut-offs | Precision, and up to an hour of staleness |
That is what requirements gathering is: finding the questions whose answers pick the design, asking them before you draw, and turning the answers into numbers you can defend. This lesson teaches six moves for doing that. Each one starts from a pressure you can spot in the prompt or in an answer.
| # | Pressure you notice | Move | What you leave with |
|---|---|---|---|
| 1 | The prompt names a whole product | Pin the core actions, say what's out | 2β3 functional requirements and an out-of-scope list |
| 2 | Two plausible answers lead to two different diagrams | Ask the forking question | The few non-functional requirements that drive the design |
| 3 | Adjectives where numbers should be ("huge", "real-time") | Turn answers into numbers | Peak QPS, storage and bandwidth, and which one binds |
| 4 | "Make whatever assumptions you like" | Assume out loud, with a number and a reason | Assumptions the interviewer can correct |
| 5 | The requirements live only in your head | Write them down, read them back, trace to them | A board you design against |
| 6 | The design can't meet a requirement | Flag the conflict and renegotiate | An explicit, agreed trade-off |
This lesson is the deep treatment of Scope, the first phase of the course's interview framework (Scope β Sketch β Deep dive β Wrap-up) from Interview Framework & Strategy. In a 45-minute slot, Scope runs from minute 3 to minute 10, right after the intro and the prompt, and closes at the minute-10 checkpoint with features, qualities, two or three numbers and your assumptions on the board. In a 60-minute slot it runs from minute 5 to 13; the extra time buys a second deep dive, and Scope grows by only a minute. Moves 1β5 happen inside Scope (Move 5's tracing then continues through the design). Move 6 happens later, when the design runs into a requirement. Design Process Steps picks up at the Sketch: API, data model and the first diagram. For latency numbers and more estimation practice, see Foundations of System Design.
Move 1: Pin the core actions and say what's out
Pressure: the prompt names a whole product. "Design Uber" covers dispatch, pricing, payments, fraud detection, maps, driver onboarding and support tools. In 45 minutes you can sketch two or three of those and design one in depth, or cover all of them badly.
A functional requirement is a behaviour: something a user or another system can do. Write each one as actor + verb + object: "A rider can request a ride." "A driver can accept a request." Each line later becomes an API call and some data, which is where the Sketch starts (Design Process Steps).
Read the prompt in three layers
| Layer | What it holds | What you do with it |
|---|---|---|
| Stated | The words of the prompt, especially adjectives: "real-time chat", "global file storage" | Treat each adjective as a hint about which quality the interviewer cares about |
| Implied | What the domain always brings | State it; don't spend a question on it |
| Missing | What only the interviewer knows: which features, how big, which qualities matter most | This is where your questions go |
Implied content is the physics of the domain. A feed implies fan-out. A payment implies that charging twice is a bug. Notifications imply retries, user opt-outs and rate limits. Anything with seats or stock implies "never sell the same item twice". You show more judgment by stating these than by asking about them.
Find the actors, then walk one workflow
Ask who or what calls the system: a rider app, a driver app, an admin console, or, for infrastructure prompts, other services. Different actors have different traffic. Drivers write their location constantly; riders mostly read it.
Then walk the most important workflow once, step by step. Hidden requirements fall out:
Rider requests a ride
1. Rider opens the app, sees nearby cars -> find drivers near a point (frequent read)
2. Rider enters a destination, sees a fare -> pricing (in scope?)
3. Rider confirms -> matching: exactly one driver per rider
4. Driver accepts; both watch the car move -> a live location stream
5. Trip ends -> payment, receipt, rating (in scope?)
Five requirements came out of that walk: a nearby-driver search, a fare estimate, one-to-one matching, a live location stream and a payment tail. Two of them, pricing and payment, came with a scoping question attached.
Say what you are not building
Pick the two or three actions the interviewer cares about most, and put the rest on an explicit out-of-scope list. Skipping payments without mentioning it doesn't count as scoping; the interviewer may think you forgot it.
What to say:
"I'll focus on requesting a ride, matching a driver and live tracking. I'll treat payments as an external service we call once at the end, and leave out onboarding, surge pricing and support tools. Does that match what you'd like to see?"
A good way to get the interviewer's priorities is to ask them to choose rather than to list: "If we keep only the three most important user actions, which are they?"
Likely follow-up: "What about pricing?" Answer with a one-line decision, not a new design: "I'll assume a pricing service returns a fare estimate. If we have time at the end, I'll come back to surge pricing."
Your turn: the prompt is "Design a notification system." Which single functional question changes the architecture the most? And which implied requirements would you state without asking?
Check your answer
Ask who triggers notifications and on which channels: push, email, SMS, in-app, or all of them? Alerts for an internal operations team on one channel are a small service. Three consumer channels through third-party providers need per-channel queues, provider rate limits and retries.
State without asking: failed sends are retried; a retried send carries an idempotency key so the worker (and the provider, where it supports one) can drop repeats, which reduces duplicates but can't rule out every one; user opt-outs are respected; and no user gets flooded (a per-user rate limit).
Move 2: Ask the questions that fork the design
Pressure: there is a question where two plausible answers lead to two different diagrams. Those are the questions worth your minutes.
The filter
Before you ask anything, check: would a different answer change a box or an arrow? If yes, ask. If no, state an assumption and move on.
- "Should it use HTTPS?" Assume it.
- "Should I use SQL or NoSQL?" That is a solution question. Choosing is your job, based on the requirements.
- "How many engineers will maintain it?" It rarely changes the diagram. Skip it.
- "How many orders arrive in the dinner hour?" Ask. The answer decides what you scale.
Non-functional requirements: the qualities that fork designs
A non-functional requirement says how well the system must behave: how much load, how fast, how available, how fresh, how durable, and where. These drive most architecture decisions. Two systems with identical features can need completely different designs because of one of these numbers.
| Quality | Ask it like this | Answers that fork the design |
|---|---|---|
| Load | "How many daily users, and what does a typical one do per day?" | Hundreds of requests/s: one database and a cache. Hundreds of thousands: replicas, partitioning, caching tiers |
| Peak shape | "What does the busiest minute look like?" | A smooth 3Γ evening peak: size for it. A ticket sale that opens at 10:00: admission control and a queue |
| Read/write mix | "Roughly how many reads per write, and which path has the latency target?" | 100:1 reads: caches and read replicas. Write-heavy: partitioned writes, append-friendly storage |
| Latency | "What p99 does the critical request need?" | Under 100 ms: precompute and cache, and keep slow work off the request path. A few seconds: more of the work can happen inside the request |
| Freshness | "After a write, who must see it, and how soon?" | "The writer, immediately; others within a minute": read-your-writes plus async replicas. "Everyone, always": one authoritative copy for that data |
| Availability | "What does an hour of downtime cost, and which features may degrade?" | 99.9% measured over a year (8.8 h): a standby you can fail over to, even by hand. The same 99.9% measured per month is only 43 min, so one slow manual failover can use it up: automate it. 99.99% (53 min a year): automatic failover, no human in the path. Surviving the loss of a whole region: multi-region |
| Durability | "Which data can we never lose, and which can we rebuild?" | "Never lose an acknowledged payment": acknowledge only after a second machine has it. "Rebuildable": a cache or a single copy is fine |
| Geography | "Where are the users, and must data stay in a country?" | One country: one region. Worldwide users with heavy media: CDN and regional copies. Residency rules: per-country storage |
| Delivery | "Is a duplicate worse than a loss? Does order matter?" | "No visible duplicates": idempotent consumers. "Per-user order": ordering by key |
| Compliance | "Is there payment, health or personal data?" | Encryption, audit logs, retention and deletion rules |
You won't ask all ten. Scope usually holds four to seven forking questions in total, and three or four of them are about the qualities this prompt makes risky. Assume the rest out loud. Key Concepts & Terminology covers availability, consistency and tail latency in depth. Databases & Storage and Messaging & Queues cover the designs these answers lead to.
Five pairs that look alike but aren't
Candidates lose points by merging these. Keep them apart in your questions and in your summary.
- Read-heavy does not mean stale reads are fine. The read/write ratio tells you where the load goes. Whether stale data is acceptable comes from the use case. A bank balance is read far more often than it is written, and it still must not be stale.
- Delivery speed is not consistency. "How quickly must an update reach users?" and "Can a reader see old data, and for how long?" are separate questions. Chat delivery must be fast, yet a message may reach your second device a moment later. A nightly ledger report is slow and must be exact.
- More nines is a redundancy problem, not a consensus problem. Availability comes from redundant copies and automatic failover (which may lean on a small consensus service to pick the new leader safely). A consensus protocol such as Raft or Paxos is what you reach for when replicas must agree on the order of writes, and during a network partition its minority side stops accepting writes, which costs availability there.
- Surviving a crashed node is not "choosing AP". A majority-quorum system keeps working when a minority of its nodes die, and it is still consistent. The CAP trade-off appears only when the network splits the nodes into groups that can't talk to each other.
- A write-ahead log is not durability against losing the machine. A WAL flushed to disk survives a crash and restart of the same node. If the disk itself is gone, only a copy on another machine helps. PostgreSQL's documentation says that with asynchronous shipping there is "a window for data loss", and that synchronous replication loses committed data only "if both the primary and the standby suffer crashes at the same time" (PostgreSQL docs).
Ask for numbers and consequences, not yes or no
"Do we need high availability?" gets "yes", because everyone wants it. You learn one bit. Ask for the number or the consequence instead:
| Yes/no opener | Question that returns information |
|---|---|
| "Do we need high availability?" | "What happens to the business if checkout is down for an hour?" |
| "Is it read-heavy?" | "Roughly how many reads per write, and are there spikes?" |
| "Should it be global?" | "Where are the users, and must any data stay in one country?" |
| "Do we need real-time updates?" | "How long after an event may a user see it, and what does a 30-second delay cost?" |
Questions that ask for a number ("How many daily users?") are exactly right. Yes/no questions are fine as follow-ups that confirm what you heard: "So a minute of staleness is OK for other people's feeds?"
Order: the biggest fork first
Ask the highest-impact unknown first. For most prompts that is scope (what are we building?), then load (how big?), then the one or two qualities the domain makes dangerous:
- money, seats, stock β correctness: state the obvious invariant ("never sold twice"), then ask what else must never happen
- chat, live tracking β latency, ordering, connections
- audio, video, photos β bytes and bandwidth
- a public API β availability and abuse
Predict first: "Design the ticketing system for a stadium concert." You know it's 60,000 seats and every buyer is in one country. Which quality requirement do you state without asking, and which question do you ask first?
Check your answer
- State it: "I'll take it that a seat can never be sold twice." Every ticketing product needs this, so don't spend a question on it; say it and let the interviewer nod. Seat inventory then needs one authoritative copy with conditional writes: "sell seat 14F only if it is still free". That is strong consistency for a small piece of data, not for the whole system.
- Ask it: "How many buyers arrive when the sale opens?" If one million people arrive in the first minute for 60,000 seats, that is about 16,700 new buyers per second, each making several requests. The design is then about admission (a waiting room that lets buyers in at a rate the checkout can handle) and holding seats for a few minutes during payment. Average daily load hardly matters.
Everything else, such as catalog size or how long to keep order history, you can assume.
Move 3: Turn answers into numbers
Pressure: adjectives where numbers should be: "lots of users", "big files", "fast". A number tells you which part of the design is under pressure and, just as often, which part isn't.
The chain you run for almost every prompt:
DAU Γ actions per user per day = requests per day
requests per day Γ· 86,400 = average per second (1M a day β 12/s)
average Γ peak factor = peak per second (size for this)
new items per day Γ bytes Γ days = storage (Γ copies you run yourself)
peak requests/s Γ bytes per response = egress, bytes/s (Γ 8 for bits/s)
Worked example: a podcast app
The interviewer has told you: 10 million daily listeners; each listens for about 45 minutes a day; the evening commute peak is about 3Γ the average; about 30,000 new episodes are published a day and kept forever, so you plan for five years; and a listener can resume on any device, losing at most 30 seconds of their place. You assume, out loud: audio at 96 kbit/s; a 40-minute average episode (about 29 MB); 2 KB of catalog metadata per episode; 10 catalog or show pages opened per listener per day. The app saves the listening position every 30 seconds while playing, which is what the 30-second resume rule allows.
DAY = 86_400 # seconds per day
# Given by the interviewer
listeners = 10_000_000 # daily active listeners
listen_s = 45 * 60 # listening per listener per day
peak = 3 # evening peak vs daily average
new_episodes = 30_000 # published per day
save_every_s = 30 # "resume within 30 s on any device"
# Assumed by you (say them out loud)
audio_bytes_per_s = 96_000 // 8 # 96 kbit/s audio = 12,000 bytes/s
episode_bytes = 40 * 60 * audio_bytes_per_s # a 40-minute episode, about 29 MB
catalog_row = 2_000 # bytes of metadata per episode
pages_per_listener = 10 # catalog and show pages opened per day
def rate(per_day):
return f"{per_day / DAY:,.0f}/s avg, {per_day / DAY * peak:,.0f}/s peak"
egress = listeners * listen_s * audio_bytes_per_s # bytes per day
print(f"audio egress {egress/DAY*8/1e9:.0f} Gbit/s avg, {egress/DAY*peak*8/1e9:.0f} Gbit/s peak, {egress/1e12:,.0f} TB/day")
audio = new_episodes * episode_bytes # bytes per day
print(f"audio stored {audio/1e9:,.0f} GB/day, {audio*365/1e12:,.0f} TB/yr, {audio*365*5/1e15:.1f} PB in 5 yr")
print("resume writes ", rate(listeners * listen_s / save_every_s))
print("catalog reads ", rate(listeners * pages_per_listener))
catalog = new_episodes * catalog_row
print(f"catalog size {catalog/1e6:.0f} MB/day, {catalog*365*5/1e9:,.0f} GB in 5 yr")
| Number | Value | What it decides |
|---|---|---|
| Audio egress | 30 Gbit/s average, 90 Gbit/s peak, 324 TB/day | A CDN in front of the audio files, which never change and so cache well |
| Audio stored | 864 GB/day, 315 TB/yr, 1.6 PB in 5 years | Object storage, not a database |
| Resume writes | 10,417/s average, 31,250/s peak | The busiest table: tiny values, overwritten constantly |
| Catalog reads | 1,157/s average, 3,472/s peak | A cache in front of one database |
| Catalog size | 60 MB/day, 110 GB in 5 years | Fits on one database node, plus replicas, for years |
Why the numbers earn their minutes
Each number points to a different decision, and two of them point away from the obvious. The catalog, which most candidates start drawing first, is the easy part: 110 GB after five years. The busiest database table is one the prompt never mentioned: resume positions, 31,250 writes/s at peak. Each write is a few dozen bytes that replaces the previous one, and only the latest position per listener and episode matters, a few GB in total. That calls for a key-value store keyed by listener and episode, not more rows in the catalog database. Notice also where the CDN comes from: egress volume and distance to listeners, not the number of bytes stored. A petabyte archive that is rarely played needs no CDN; a small set of episodes streamed by millions of listeners does.
In the interview you don't recite the whole table. Scope keeps only the two or three numbers that decide something, here 90 Gbit/s of egress, 1.6 PB of audio and 31,250 resume writes/s, each with the decision it drives. Inputs the interviewer gave, such as the 10 million listeners, stay on the board but aren't deciding numbers. The rest stays in your notes.
π§ Say it like this: "Our load is bytes plus one small, hot table. Audio is 90 Gbit/s at peak and 1.6 PB over five years, so it goes into object storage behind a CDN. The catalog is only about 110 GB, so one database with a cache is enough. The surprise is resume positions: about 31,000 writes a second of tiny values that overwrite each other, so they get their own key-value store."
Likely follow-up: "Can resume be exact to the second?" Recompute; don't guess. The next exercise does exactly that.
Unit traps
| Trap | Example | Fix |
|---|---|---|
| Bytes vs bits | 11.25 GB/s looks as if it fits on two 10 Gbit/s links | Network links are rated in bits: 11.25 GB/s Γ 8 = 90 Gbit/s |
| Per day vs per second | "50M requests" per day or per second? | Divide by 86,400. For mental maths, 1M a day β 12/s |
| Average vs peak | Sizing resume writes for 10,417/s when the evening brings 31,250/s | Always ask for the peak shape. 3Γ suits a smooth evening peak; a sale opening can be 100Γ |
| TB vs TiB | 1 TB = 10ΒΉΒ² bytes = 0.91 TiB | Pick one unit, label it, and don't mix them in one sum |
| Copies | Sizing a database for one copy when you run three replicas | Multiply by the copies you run yourself. A managed object store bills logical bytes and keeps its own copies |
Round honestly and out loud: "about 31,000 writes a second", not "31,250".
Pick the right unit of load
Requests per second is not always the number that matters. Chat is measured in messages per second and open connections. Live tracking is measured in location updates per second. Audio and video streaming are measured in bytes per second. Metadata stores are measured in rows. In the ride-sharing walkthrough below, the hard number turns out to be location updates, not trips.
Your turn: the product manager asks for resume "exact to within 5 seconds, on any device". If the app simply saves the position every 5 seconds instead of every 30, what is the peak resume write rate? What would you say back?
Check your answer
Saving six times as often multiplies the writes by six: 31,250 Γ 6 = 187,500 writes/s at peak, each for a value that is overwritten 5 seconds later.
Ask what the rule is really for. It is switching devices, and a listener almost always pauses or closes the app before switching. So save on every pause, stop and app close, which meets "within 5 seconds" in normal use, and keep a slower periodic save, say every 2 minutes, as crash recovery. The periodic part is then about 7,800 writes/s at peak, a quarter of today's, plus a few pause events per session. Then state the price out loud: "If the app crashes, a listener can lose up to 2 minutes. Exact to 5 seconds even after a crash costs six times today's writes. Which do you want?" Requirements have prices, and naming them is the move.
Move 4: Assume out loud when the interviewer says "you decide"
Pressure: "Make whatever assumptions you like." This isn't a trap. It tests whether you can drive the conversation.
State the assumption with a number, a reason and an invitation to correct it:
"I'll assume 50 million daily users, each reading about ten items for every one they write, like a large consumer app. If this is an internal tool, the design gets much simpler, and I'll point out where."
Four rules keep assumptions honest:
- Don't dress an assumption up as a fact. "I'll assume links don't expire unless the creator sets an expiry" is a fine assumption. "β¦because 30 days is the industry default" is a claim you can't back up, and the interviewer may know it's false.
- Label every number on the board as given, assumed or derived. Never present a number you made up as something the interviewer said. If they said "mostly reads", you have a ratio, not a request rate.
- Ask when the answer forks the design and you can't guess it. Assume when the domain answers it or it doesn't change the diagram.
- If a high-impact unknown stays open, choose the branch that makes the more interesting design, and say what the other branch would change: "I'll design for worldwide users. If it's one country, I'd drop the multi-region part."
Too many questions, or too few
Over-questioning looks like a dozen questions, several of which you should answer yourself ("SQL or NoSQL?") and several of which change nothing. The interviewer concludes that you can't make judgment calls. Under-questioning looks like one question or none, with silent assumptions everywhere. Nothing seems wrong until the mismatch surfaces twenty minutes in. Here is what silent assumptions produce for "Design a messaging system":
| Candidate A | Candidate B | |
|---|---|---|
| Assumed users | Millions, worldwide | 10,000 inside one company |
| Assumed content | Text, media, reactions | Text only |
| Assumed history | Delivered, then rarely read again | Permanent, searchable archive |
| Assumed availability | 99.99%, consumer app | 99.9%, business hours matter most |
| Resulting design | Partitioned, multi-region, media behind a CDN | One application, PostgreSQL, object storage for attachments |
Both designs are coherent, and at most one of them matches what the interviewer had in mind. A healthy Scope phase has a handful of forking questions (often four to seven), stated assumptions for everything else, and a read-back.
Your turn: for "Design a food-delivery app", sort each question into ask, assume aloud or decide yourself:
- "Should the API use HTTPS?"
- "Should orders go in PostgreSQL or Cassandra?"
- "How many orders a day, and how sharp is the dinner peak?"
- "Can one courier carry two orders at once?"
- "Do customers need an account to order?"
- "How fresh must the courier's position be on the customer's map?"
Check your answer
- Ask: 3 and 4. Load and peak shape decide what you scale. Batching two orders per courier turns matching into a routing problem.
- Ask, or assume if you are short on time: 6. "Every few seconds" is a safe assumption to state; the interviewer will correct it if they care.
- Assume aloud: 1 (yes) and 5 (yes, accounts already exist).
- Decide yourself: 2. It is a solution question, and you will answer it once you know 3.
Move 5: Write it down, read it back, trace every decision
Pressure: the requirements live only in your head, so neither you nor the interviewer can point at them later.
Write them where both of you can see them, with short ids and a source for every number. Here is the board for the podcast app from Move 3, after two more questions added a playback-start target and a durability rule:
FUNCTIONAL
F1 A listener can play an episode and resume it on any device
F2 A listener can follow shows and see their new episodes
F3 A publisher can upload an episode
NON-FUNCTIONAL source
N1 10M daily listeners x 45 min; 30k new episodes/day given
N2 Peak 3x average: 90 Gbit/s audio, 31k resume writes/s given + assumed + derived
N3 Playback starts within 2 s at p99 given
N4 A published episode is never lost given
N5 Resume within 30 s of the listener's place, on any device given
N6 Keep episodes forever; 5 yr = 1.6 PB audio, 110 GB catalog given + assumed + derived
OUT OF SCOPE
Recommendations, search, ads, comments, live shows
Three figures on this board decide something: 90 Gbit/s, 1.6 PB and 31k resume writes/s. The rest are inputs to them or bounds the design must meet. The ids pay off during the design. Every decision becomes a sentence that points at a line:
- "Because N2 is 90 Gbit/s at peak and N6 is 1.6 PB, audio goes into object storage behind a CDN."
- "N5 lets a listener lose up to 30 seconds of their place, so the app saves every 30 seconds and on pause, and the store keeps only the latest position per listener and episode."
When the interviewer asks "Why not keep the audio in the database?", you point at N6 instead of defending a preference.
Requirements, not solutions
A requirement says what must be true. A solution says how you'll make it true. "We need Kafka" and "use WebSockets" belong in the Sketch, after Scope. Here is a test that works: delete every product, protocol, data structure and algorithm name from the line. If nothing is left, it was a solution. If something is left, that remainder is the requirement.
| Line on the board | After deleting the mechanism | The requirement to write |
|---|---|---|
| "Use Kafka so events survive consumer outages" | "events survive consumer outages" | "No event is lost if a consumer is down for up to 6 hours" |
| "WAL required" | (nothing) | "An acknowledged write survives the loss of any one machine" (a WAL alone can't meet this) |
| "Use WebSockets for driver location" | (nothing) | "The rider's map shows the driver's position at most 5 s old" |
| "Put a CDN in front" | (nothing) | "Images load in under 300 ms for users in Europe and Asia" |
The requirement version does two jobs. It keeps the choice open: polling can also meet "at most 5 s old", as long as the driver's ping interval plus the poll interval plus network time stays under 5 seconds (with 4-second pings, poll about twice a second). And it tells you how to check whichever solution you pick: retention must outlast the 6-hour outage, and polling every 4 seconds would show positions up to about 8 seconds old. The WAL row shows the same payoff. Once the need is written down, it's plain that a log on one machine's disk can't meet it; the write has to reach a second machine before you acknowledge it.
Read it back
Before the first box, spend 30 seconds on it: "Here's what I'll design against: β¦ Did I miss anything you care about?" A correction now costs a sentence. The same correction at minute 25 costs the design.
When to stop asking: Scope closes at the minute-10 checkpoint. By then the board holds what Scope produces: the features (with what's out), the qualities this prompt makes risky, the two or three numbers that decide something, and your assumptions. Read it back and move to the Sketch. If you're running late, compress rather than skip: fewer questions, more stated assumptions, and still the read-back. Anything else can be settled when the design reaches it.
Move 6: Flag the conflict when a requirement meets the design
Pressure: twenty minutes in, your design can't meet a requirement, or the interviewer adds a twist. There are two bad responses: quietly break the requirement and hope nobody notices, or panic and start over. The good response is to flag the conflict: name it, offer options with their costs, and let the interviewer choose. This move happens after Scope, in the Sketch or the Deep dive, and it works only because Scope left a board to point at.
Here is a conflict that physics creates. A product-page API promises p99 under 50 ms in every region, and its one database sits in Virginia. Singapore is about 15,500 km from Virginia by great circle. Light in optical fibre covers about 200,000 km/s, so one round trip needs at least 155 ms before any router, queue or TLS handshake adds to it. A p99 target ignores the slowest 1% of requests, so a few slow misses would be fine. But about 5% of Singapore's requests miss the cache, which puts misses inside Singapore's p99: its p99 is at least 155 ms.
"I want to flag a conflict. We agreed p99 under 50 ms in every region, but about 5% of Singapore's requests miss the cache, and each miss to our database in Virginia costs at least 155 ms of round trip. That's the speed of light, so no database tuning fixes it. Option one: put a read replica in Asia, which costs replication and a few seconds of lag for new products. Option two: push Singapore's hit rate above 99% by pre-warming the popular pages, so misses fall inside the 1% that p99 ignores. Option three: agree that Singapore's p99 is about 250 ms while its p90 stays under 50 ms. Which fits the product better?"
That one move shows the interviewer three things. You track requirements while you design, you know what they cost, and you negotiate instead of ignoring them.
Flags also run the other way. When the interviewer adds information ("What if 30% of users are in Asia?"), update the board line, re-derive the numbers it feeds, and say what shifts: "With 30% of traffic in Asia and 5% of it missing, Asian misses alone are 1.5% of all requests, so even a global p99 would fail. Let me add a read copy in Asia." (Europe needs the same check: London to Virginia is at least a 59 ms round trip.)
Walkthrough: a ride-sharing service, one question at a time
"Design a ride-sharing service." Here is the Scope phase as a trace. Each row is one question, the interviewer's answer, and what that answer changes.
| # | You ask | Interviewer says | What it changes | Board |
|---|---|---|---|---|
| 1 | "Which part of a ride should we focus on: requesting and matching, live tracking, pricing, payments?" | "Requesting, matching and live tracking. Payments are a black box." | Three flows; the rest is out of scope | F1βF3, OUT |
| 2 | "How many trips a day, and how many drivers are online at the busiest hour?" | "About 20M trips a day. At peak, 1M drivers online. Evenings run about 3Γ average." | The rates below | N1 |
| 3 | "How stale can the car on the rider's map be? How fast must a match appear?" | "Drivers send a position every 4 s. A match within 10 s." | Location is soft real time; matching takes seconds, not milliseconds | N2, N3 |
| 4 | "I'll take it that a driver is never assigned to two riders at once. What else must never happen?" | "Right. And a rider is never charged for a ride that didn't happen." | Strong consistency for the assignment only; a charge only after a completed trip | N4, N8 |
| 5 | "Where do we operate, and does a trip ever cross regions?" | "Many countries. A trip stays in one city." | The city is a natural partition key | N5 |
| 6 | "Which features must keep working during an outage, and what's the budget?" | "Requesting and in-progress trips: 99.99%. Trip history can be down for a while." | Automatic failover; availability set per feature | N6 |
| 7 | "Do we keep location history?" | "Only each trip's route, for receipts and disputes." | The latest position is disposable; routes are durable | N7 |
Seven short questions, each of which forks something; with the arithmetic and the read-back they fill Scope's seven minutes. Of the numbers that follow, one decides the design (250,000 location writes/s); the trip rate is on the board to show it is small. Now turn rows 2 and 3 into numbers. The trip length and the size of one location ping are your assumptions:
DAY = 86_400
trips_per_day = 20_000_000 # given
peak = 3 # given: evenings are about 3x average
drivers_online_peak = 1_000_000 # given
ping_every_s = 4 # given
trip_s = 15 * 60 # assumed average trip length
ping_bytes = 100 # assumed: id, lat, lon, time, heading
requests_peak = trips_per_day / DAY * peak
writes = drivers_online_peak / ping_every_s
active_trips = requests_peak * trip_s # Little's law: rate x time in system
print(f"trip requests: {trips_per_day/DAY:,.0f}/s avg, {requests_peak:,.0f}/s peak")
print(f"location writes: {writes:,.0f}/s at peak")
print(f"trips in progress at peak: {active_trips:,.0f} -> {active_trips/ping_every_s:,.0f} map updates/s")
print(f"all latest positions: {drivers_online_peak*ping_bytes/1e6:,.0f} MB")
print(f"trip routes: {trips_per_day*trip_s/ping_every_s*ping_bytes/1e9:,.0f} GB/day")
| Output | Value |
|---|---|
| Trip requests | 231/s average, 694/s peak |
| Location writes | 250,000/s at peak |
| Trips in progress at peak | 625,000, so about 156,000 map updates/s pushed to riders |
| All latest positions | 100 MB |
| Trip routes | 450 GB/day |
Predict first: which number drives the design, and what does it rule out?
Check your answer
250,000 location writes per second of data that is worthless four seconds later. Trips are only 694/s at peak, which is small.
Writing each ping as a durable, replicated row in the trips database would pay for disk flushes and replication on a value that is about to be overwritten. A partitioned relational database could absorb the load; the problem is that nearly all of that work is wasted. So keep the latest positions in memory, partitioned by city, in an index that answers "drivers near this point". All of them together are only 100 MB. Append each trip's route to durable storage in batches.
What you give up: when a location node crashes, it forgets its positions. Every driver resends within 4 seconds, so that state rebuilds itself. Say that trade-off out loud; it is the sentence the interviewer is waiting for.
The other rows shape the rest of the design:
- Row 4 β one small piece of strong consistency. Matching needs a conditional write: "assign driver D to rider R only if D is still available". If two matchers race for the same driver, one of them loses and tries the next driver. The charge rule is simpler: the trip service asks the payment service for a charge only when the trip reaches "completed". Positions and arrival estimates can be approximate.
- Row 5 β partition by city. Almost every query is local, and a failing region takes down only the cities it serves.
- Row 6 β availability per feature. 99.99% allows about 4.4 minutes of downtime a month. One manual failover would use it up, so failover must be automatic. Trip history gets a looser target and a simpler setup.
=== Ride-sharing requirements ===
F1 A rider can request a ride and see a fare estimate (pricing is a black box)
F2 The system matches the request to one nearby available driver
F3 Rider and driver both see the car's live position until drop-off
N1 20M trips/day; peak 3x -> 694 requests/s given + derived
1M drivers online at peak -> 250k location writes/s given + derived
N2 Position on the map at most about 5 s old derived (4 s pings)
N3 Match within 10 s of the request given
N4 A driver is never assigned to two riders at once assumed, confirmed
N5 Trips never cross cities; many countries given
N6 Request + in-progress trips 99.99%; history may degrade given
N7 Keep each trip's route; the latest position is disposable given
N8 A charge is requested only for a completed trip given
OUT Payments internals, onboarding, surge pricing, support tools
What breaks first when the business grows 10Γ? Location writes reach 2.5 million per second. Partitioning by city spreads most of the load, but the biggest cities become hot partitions, so they have to be split further into smaller geographic cells. Trip requests, at about 7,000/s, are still not the problem.
Walkthrough: a key-value store, where one answer flips the design
"Design a distributed key-value store." This is an infrastructure prompt. Its users are other services, and the first question is different: "What will teams store in it?" Watch two answers take the same prompt to opposite designs.
| Answer A: "Feature flags and service configuration" | Answer B: "User profiles and preferences for our app" | |
|---|---|---|
| Data size | 20,000 keys Γ 5 KB β 100 MB | 500M profiles Γ 2 KB β 1 TB |
| Writes | A few per minute | 20,000/s at peak (settings, last-seen times) |
| Reads | Services cache locally and watch for changes | 100,000/s at peak |
| Freshness | Every service sees changes in the same order; a flag never flips back | A few seconds of staleness is fine |
| Failure | Keep serving with a minority of nodes down | Lose any one node without losing an acknowledged write |
| Design | 3 or 5 nodes running a consensus protocol (Raft in etcd and Consul, Zab in ZooKeeper), every node holding all the data | Data partitioned by key, 3 replicas per partition, any of which accepts writes (leaderless, the Dynamo and Cassandra style); a write is acknowledged once 2 of the 3 have it |
| What you give up | Every write needs a majority; during a network partition the minority side can't accept writes | Readers can see stale values. Concurrent writes to one key must be reconciled, and last-write-wins silently keeps only one of them; if every write matters, keep all versions for the app to merge, or use conditional writes |
Answer A is small, critical data. etcd's documentation sets a default storage quota of 2 GiB and suggests 8 GiB as the maximum for normal environments (etcd limits). These systems are built for small data that everyone must agree on, not for terabytes. Answer B is a large, partitioned store where throughput sets the node count. Here is its sizing, with the per-node throughput stated as assumptions:
import math
data_tb, replicas = 1.0, 3
usable_tb_per_node = 1.0 # 2 TB disk kept half full (compaction, growth)
reads, writes = 100_000, 20_000 # per second at peak (given)
reads_per_node = 20_000 # assumed; benchmark your own hardware
replica_writes_per_node = 25_000 # assumed; every write lands on 3 replicas
need = {
"replicas": replicas,
"storage": math.ceil(data_tb * replicas / usable_tb_per_node),
"reads": math.ceil(reads / reads_per_node),
"writes": math.ceil(writes * replicas / replica_writes_per_node),
}
binding = max(need, key=need.get)
print(need)
print(f"{need[binding]} nodes, bound by {binding}; add 1 so a node can die at peak")
{'replicas': 3, 'storage': 3, 'reads': 5, 'writes': 3}
5 nodes, bound by reads; add 1 so a node can die at peak
Six nodes, and read throughput sets the count under these assumptions. If one node served 50,000 reads/s, storage, write throughput and the replica count would all bind at three: four nodes with the spare. Say which assumption the answer hangs on.
Two failure requirements deserve a precise sentence each:
- "Lose any one node without losing an acknowledged write." Acknowledge a write only after two of the three replicas have it. A local write-ahead log alone doesn't meet this requirement: it survives a restart, not a dead disk. Neither does a leader that replicates asynchronously, because it can acknowledge a write and then die before any follower has it.
- "Keep serving when a node dies." Both designs do this. A 3-node Raft cluster keeps its majority with one node down, and a 2-of-3 write still finds two live replicas. The designs differ only when the network splits the nodes into groups. Answer A stops writes on the minority side, always. Answer B can choose per request: requiring 2 of 3 replicas refuses writes where only one replica is reachable, while accepting a single replica keeps writes flowing but breaks the node-loss durability requirement. That is a partition-time trade-off between accepting writes and keeping them durable. It has the same shape as CAP, but the property at stake is durability, not linearizability; Answer B isn't linearizable anyway, since its readers can see stale values. Either way, the choice appears only when the network splits, not when nodes crash.
Traps that cost candidates the interview
| Trap | What the interviewer sees | The fix |
|---|---|---|
| Designing in the first minute | A correct answer to a question nobody asked | Scope first: core actions, out of scope, then numbers |
| Interrogating: a dozen questions | Someone who can't make a call | Apply the filter; assume the rest aloud |
| Asking nothing | Silent assumptions that surface at minute 20 | Ask the four to seven forking questions |
| Yes/no openers | "Yes" to everything, and nothing learned | Ask for the number or the consequence |
| Solutions on the requirements board | Options closed before they were weighed | Delete the product name; keep what remains |
| Invented numbers presented as given | A design resting on numbers nobody agreed to | Label each number given, assumed or derived |
| Everything critical: 99.999%, strongly consistent, sub-10 ms | No sense of cost | Ask what failure costs; requirements have prices |
| Requirements frozen once Scope closes | Violations hidden in the design | Flag the conflict with options and costs |
Final round: no label on the prompt
Real prompts don't say which move they need. Start with scope, find the forks, then get the numbers.
Challenge 1: a price-alert service
"Users set a target price on a product; we email them when the price drops below it." You have time for four questions. Which four, in what order, and what do you assume aloud?
Check your answer
- Scope: "Where do price changes come from: our own catalog, or scraping other stores? Email only, or push too?" Scraping adds a whole crawler system, so this is the biggest fork.
- Load: "How many active alerts, how many products, and how many price changes a day?" For example, 50M alerts on 5M products with 20M price changes a day is 231 changes/s. At 10 alerts per product that is about 2,300 alert checks/s, which is modest if alerts are indexed by product.
- Freshness: "How soon after a drop must the email go out?" "Within minutes" allows batching. "Within seconds" needs a streaming path.
- Duplicates: "If the price drops, rises and drops again, does the user get two emails?" This decides whether you store per-alert state ("already notified").
Assume aloud: users can set many alerts, email is sent through a provider with its own rate limits, and unsubscribe is handled.
Challenge 2: a metrics collection system
"Collect metrics from our 40,000 servers." The interviewer adds: each server emits 300 metrics every 10 seconds; raw points are kept for 15 days, then 1-minute rollups for a year. Assume 16 bytes per raw point uncompressed (a timestamp and a value) and 40 bytes per rollup point (min, max, sum, count, timestamp). Which number surprises you?
Check your answer
- Ingest: 40,000 Γ 300 Γ· 10 = 1.2 million points/s, about 19 MB/s. That is 1.66 TB a day raw, about 25 TB for 15 days.
- Rollups: 12 million series Γ 1,440 minutes Γ 40 bytes = 691 GB a day, about 252 TB for a year.
The surprise is that the "small" long-term rollups take ten times the storage of the raw data. That is a requirements conversation to have now: "Would 5-minute rollups be enough for the year?" That change alone cuts the year to about 50 TB. The write path, 1.2 million points a second, needs batching, and most points are never read one by one, which favours an append-optimised time-series store.
Challenge 3: defend the simple design
"An internal tool for 3,000 employees to book meeting rooms. No double bookings." You propose one PostgreSQL database with a standby. The interviewer pushes back: "Shouldn't we shard it and put Kafka in front?" Answer in two or three sentences.
Check your answer
"3,000 people making about ten bookings or lookups a day is 30,000 requests a day, well under one per second on average and a few per second at the 9 a.m. peak. One database handles that with room to spare. The requirement that actually matters is no double booking, and a unique constraint on room and time slot (or PostgreSQL's exclusion constraint, if bookings are free-form time ranges) enforces it in one place. Sharding or a queue would add ways to break that guarantee and buy us nothing. If the tool grew to a million users, I'd revisit it."
Requirements justify simple designs as firmly as complex ones. Showing that the numbers don't call for a component is a strong interview move.
Cheat sheet: pressure β question β move β cost
| Pressure | Question that exposes it | Move | What you give up |
|---|---|---|---|
| The prompt names a whole product | "Which two or three actions matter most?" | Scope, plus an explicit out-of-scope list | Breadth |
| Adjectives, no numbers | "How many users, doing what, how often? What's the busiest minute?" | DAU β QPS β peak β storage β bandwidth | Two minutes of arithmetic |
| Ranks, counts, percentages | "Exact, or is 'top 3%' enough? How fresh?" | Approximate or batch when allowed | Precision, freshness |
| Money, seats, stock | State "never sold twice", then ask "What else must never happen?" | Strong consistency on that data only | Write latency; availability during a partition |
| "Users see their own post" | "Who must see a write, and how soon?" | Read-your-writes; async replicas for everyone else | Routing complexity |
| "Never lose it" | "Which data can never be lost, and which can we rebuild?" | Acknowledge after a second machine has it | A replica round trip per write |
| An uptime target | "What does an hour of downtime cost? What may degrade?" | Automatic failover; multi-region only if a region's loss must be survived | Money and operational complexity |
| Far-away users, heavy media | "Where are the users? How big are the responses?" | CDN, regional copies | Invalidation, replication lag |
| Bursts | "What happens at 10:00 when the sale opens?" | Admission control, queues, size for the peak | Waiting time during bursts |
| "You decide" | None: you state it | Assumption + number + reason + invitation | Possible correction, which is the point |
| Design meets a limit | None: you raise it | Flag it: conflict, options, costs | A few seconds; you gain trust |
Before moving on, take a prompt you haven't practised, such as "Design a food-delivery app" or "Design a hotel booking system". Set a seven-minute timer, the length of Scope in a 45-minute slot, and do this aloud: name the core actions and what's out, ask four to seven forking questions and say why each one forks, invent the interviewer's answers and turn them into peak QPS, storage and bandwidth, name the number that binds, and read the board back. If a friend can stop you at any line and ask "why?", and you can point at a requirement, you're ready for the Sketch.
Next: Design Process Steps turns this board into the Sketch and the Deep dive. Interview Framework & Strategy covers pacing the whole interview.