Interview Framework & Strategy
Develop a structured approach to tackle any system design interview question confidently.
SPACED REPETITION Β· 15 practice questions
Make this lesson stick.
Try 3 questions now. No account needed. Sample answers aren't saved.
or sign in to practice all 15The first thing that breaks is the clock
The interviewer says: "Design the backend for a ride-hailing app." You know this one, so you start drawing:
rider app, driver app βββΊ load balancer βββΊ API servers βββΊ Postgres primary βββΊ replica
β
ββββββββββΊ Redis cache
Riders, drivers, trips, payments, a cache for hot reads. Every box is defensible, and explaining them well takes twenty minutes.
At minute 26 the interviewer asks, "Where do driver locations go?" Into Postgres, you say. "How often does a driver send one?" Every few seconds, so you settle on 4. "How many drivers are online at peak?" You agree on 500,000.
That is 500,000 Γ· 4 = 125,000 location writes per second, each one also updating a spatial index, all aimed at the one primary you drew. Whether a heavily tuned box could survive that is debatable. What is not debatable is that this was the hard part of the system all along: a firehose of positions that are worthless for matching four seconds after they arrive. You found it at minute 26. The slot ends at 45 and the last five minutes are for your questions, so you have 14 minutes to redesign the core of the system.
Nothing you drew was wrong. The order was wrong. "How often do drivers report their location, and how many are online?" costs thirty seconds at minute 3.
Predict first: once you know about the 125,000 writes per second, what do you do differently with driver positions?
Check your answer
Keep only each driver's latest position, in memory, split by city, and stop writing every ping as a durable row in the database. For matching, a position is replaced four seconds later anyway, so losing one costs at most one ping interval. Pings from a driver who is on a trip are that trip's route, which fares and disputes need, so append those to a cheap log or object storage, off the hot path. Trips, payments and profiles still belong in Postgres. The rest of the diagram survives. One box changes, and it is the box the whole interview should have been about.
The clock is a constraint like any other. A 45-minute slot is not 45 minutes of design: introductions and the prompt take a few minutes, and many interviewers keep the last five for your questions. You have about 37 minutes to scope, design, go deep and wrap up. This lesson is about spending them well, and it teaches six moves. Each starts with a pressure you will feel in the room.
| # | Pressure you feel | Move |
|---|---|---|
| 1 | A one-line prompt and a blank board | Scope before you solve: ask what changes the design, write the rest down as assumptions |
| 2 | Numbers everywhere and no time | Estimate only what decides something |
| 3 | 37 minutes and four things to cover | Budget the clock with three checkpoints |
| 4 | The urge to perfect one box | Breadth first, then depth where the risk is |
| 5 | The interviewer interrupts | Read the interruption and answer in the trade-off shape |
| 6 | You are stuck, or you just spotted your own bug | Recover out loud |
This lesson assumes you know the building blocks (load balancers, caches, databases, queues) from Foundations of System Design and Core Building Blocks. It is about using them under a clock, with someone watching. The next two lessons go deeper on the phases themselves: Requirements Gathering and Design Process Steps.
Before any move: know what is being scored
The prompt is vague on purpose. Real projects start vague too, and the interviewer is simulating a design discussion with a future colleague. They want to see whether you shape the problem or wait for someone to hand you a spec.
Companies word their rubrics differently, and few publish them in detail. Interviewers' notes still tend to land in four areas:
| Area | What earns it | What loses it |
|---|---|---|
| Problem navigation | Finding the hard part early, scoping, choosing what deserves time | Designing before you know what matters; 20 minutes on the easy part |
| Solution design | A design that meets the stated requirements end to end, each choice tied to one | Boxes with no reason; a requirement that nothing in the diagram meets |
| Technical depth | Going deep where the risk is, with facts that are correct | Name-dropping ("add Kafka"); confident wrong claims |
| Communication | Reasoning they can follow; checking in; using their hints | Long silence followed by a conclusion; a monologue that ignores their questions |
Knowledge is not the minor part here. Structure makes your knowledge visible; it does not replace it. A well-organized answer built on a wrong fact still fails the depth row, and Move 5 shows one. What structure buys you is that the interviewer sees the knowledge you do have, in an order that shows judgment.
Two consequences follow:
- There is no single right design. Two candidates can draw different systems for the same prompt and both do well, if each choice is tied to a requirement and its cost is named.
- The reasoning is the answer. "I'd use a queue here" earns almost nothing on its own. The interviewer cannot tell whether you understand queues or remember the word. The same conclusion, with the requirement that forced it and the price you pay, scores on the design and depth rows.
The bar moves with the level
Many companies skip the system design round for entry-level roles, or run a lighter version. Where the round exists, the same prompt can serve several levels. What changes is how much you drive and how deep you go. The table shows how interviewers commonly calibrate; your recruiter can tell you what your target level emphasizes.
| Level | What usually passes |
|---|---|
| Mid-level | A working end-to-end design with sensible components; handles probes well, even when the interviewer points to where they should go |
| Senior | Drives the session; picks the deep dive by risk and says why; raises failure modes and trade-offs before being asked; puts numbers on claims |
| Staff and above | All of that, plus framing the ambiguity itself: comparing whole approaches, cost and operability, how the system evolves or migrates, where team boundaries fall |
β οΈ "Senior" does not mean "designs for a billion users from minute one." It means designing for the stated scale and knowing what changes at 10Γ. Over-building is a judgment error at every level (Move 4).
Move 1: Scope before you solve
Smell: you are drawing boxes, and nobody has said a number yet.
Scoping produces four things, written where the interviewer can see them:
- Core features: the three to five things the system must do. Everything else goes on an "out of scope" line.
- The qualities that matter: latency, consistency, availability, durability. Pick the two or three that bite for this prompt.
- The scale: the two or three numbers that will decide something (Move 2).
- Assumptions: everything you decided instead of asked.
Ask the questions that change the drawing
A clarifying question is worth its time only if the answer changes what you draw. "Which cloud provider?" changes nothing. "Can a rider see a driver's position a few seconds late?" changes what storage and consistency you need.
Predict first: the prompt is "Design a ticketing site for concert on-sales." Which three of these questions change the design most?
- Which payment provider do we use?
- Can a seat ever be sold to two people?
- How many fans arrive when a big show goes on sale, and how many seats are there?
- Is the frontend React or Vue?
- Are seats assigned, or is it general admission?
- Do users have profile photos?
Check your answer
2, 3 and 5. "Never sell a seat twice" means seat inventory needs atomic, strongly consistent updates, and that is the heart of the design. The size of the rush (say 200,000 fans for 20,000 seats in the first minutes, ten buyers per seat) decides whether you need a waiting room that admits buyers only as fast as the inventory can take them. Assigned seats mean holding one specific seat for a few minutes during checkout; general admission is a counter you decrement.
The payment provider matters later, when you make charges safe to retry, but it does not change the first diagram. The frontend framework and profile photos do not touch it at all.
Know when to stop
Scope gets about seven minutes of a 45-minute slot, from minute 3 to minute 10, and you do not need every answer. When the interviewer says "you decide," decide out loud: "I'll assume 500,000 drivers online at peak, each sending a position every 4 seconds. Stop me if that's off." A stated assumption costs five seconds and keeps the interview moving. An unstated one is a trap you set for yourself.
β οΈ There are two ways to get scope wrong. Zero questions gives you the opening story of this lesson. Fifteen minutes of questions leaves no time for depth. Scope is not a checklist to complete. It is a search for the requirement that will hurt.
Say the route out loud
Once the features and qualities are on the board, announce the plan:
"Here's how I'd like to use the time: a couple of numbers, then an end-to-end design with the API and data model. Then I'd go deep on location ingestion, because that's where the load is, and I'll save a few minutes for failure cases. Does that work for you?"
That sentence does three jobs. It shows structure before you have drawn anything. It gives the interviewer a cheap moment to redirect you ("I'm more interested in matching, actually"). And it gives you permission later to say, "I'll move on so we keep time for the deep dive."
For which questions to ask for different kinds of systems, functional versus non-functional requirements, and turning the answers into numbers, see Requirements Gathering.
Move 2: Estimate only what decides something
Smell: you are computing a number and you don't know what you will do with it.
Every estimate should end with "β¦so." Here are the ride-hailing numbers with their assumptions stated:
| Estimate | Arithmetic | Result | So⦠|
|---|---|---|---|
| Position writes | 500,000 drivers Γ· 4 s | 125,000 per second | Don't store each ping as a durable database row; keep the latest position per driver in memory |
| Latest positions | 500,000 Γ 100 bytes | 50 MB | Memory is not the problem; split by city for write load and blast radius, not for size |
| Trip records | 10 million trips a day Γ 1 KB | 10 GB a day, 3.65 TB a year | Partition trips by time and archive old ones; that can wait until after the first diagram |
| Driver profiles | 5 million Γ 2 KB | 10 GB | One ordinary database; spend no more time on it |
The last row matters as much as the first. Some estimates are worth zero minutes. Profile storage is far too small to change anything, so a strong candidate says "profiles are about ten gigabytes, one database" and moves on.
Three habits keep estimation fast:
- Round hard. A day is 86,400 seconds; call it 10β΅. One million requests a day is about 12 per second.
- State the assumptions before the arithmetic. Then the interviewer can correct an input instead of arguing with your result.
- Stop at the order of magnitude that decides. Whether it is 110,000 or 125,000 writes per second does not matter. That it is about 10β΅ and not 10Β³ does.
If the interviewer says "skip the estimates": skip them. They are telling you what they want to spend time on. But keep the one number that decides the design, in one sentence, as an assumption: "Sure. I'll assume on the order of 100,000 position updates a second, which is why positions live in memory." You have skipped the ceremony, not the reason.
The arithmetic itself (latency numbers, requests per second, storage and bandwidth) is covered in Foundations of System Design and Requirements Gathering.
Your turn: the prompt is a leaderboard for a mobile game. Twenty million people play each day, each finishes about 10 games, and the peak is 3Γ the average. You plan to keep every player's score in one in-memory sorted set, such as a Redis sorted set, at about 100 bytes per player. Compute the peak rate of score updates and the memory. Does either number force you to shard?
Check your answer
20,000,000 Γ 10 Γ· 86,400 β 2,300 updates per second on average, so about 7,000 per second at peak. Memory is 20 million Γ 100 bytes = 2 GB.
Neither forces sharding. A single Redis node typically handles tens of thousands of simple operations per second, and 2 GB fits in one machine's memory. So the "soβ¦" is: one sorted set, plus a replica for failover. The deep dive belongs somewhere else, such as friends-only leaderboards or resetting a weekly board. Saying this in thirty seconds shows you won't shard just to look busy.
Move 3: Budget the clock
Smell: it is minute 20, you are still on requirements, and it felt like eight minutes.
Here is the whole interview as four phases with one time budget. Use it as your default for every prompt.
| Phase | What you produce | 45-minute slot | 60-minute slot |
|---|---|---|---|
| Intro and prompt | Nothing yet | 0β3 | 0β5 |
| 1. Scope | Features, qualities, two or three numbers, assumptions | 3β10 | 5β13 |
| 2. Sketch | API, data model, and a diagram with one write and one read traced end to end | 10β20 | 13β25 |
| 3. Deep dive | The riskiest part designed properly, with its failure modes (one topic in 45 minutes, two in 60) | 20β35 | 25β48 |
| 4. Wrap-up | Summary, top risks, what changes at 10Γ | 35β40 | 48β55 |
| Your questions | Your questions for the interviewer | 40β45 | 55β60 |
In a 45-minute slot that is 37 minutes of design: 7 + 10 + 15 + 5. In 60 minutes it is 50: 8 + 12 + 23 + 7. Look where the extra 13 minutes went: 8 of them to the deep dive. A longer slot buys depth, usually a second deep-dive topic, not a longer requirements phase.
The deep dive gets the biggest share because that is where technical depth is scored and where levels separate. Everything before it exists to make sure you dive into the right thing.
Each activity has one home:
- Estimation is part of Scope: only the numbers that decide something.
- API and data model are part of the Sketch.
- Trade-offs are not a phase at the end. You state one at every decision, and the wrap-up collects the biggest ones.
- Failure modes appear in the deep dive, for the part you are deepening, and again in the wrap-up for the whole system.
Three checkpoints
For a 45-minute slot, remember three minute marks: 10, 20, 35.
- Minute 10: requirements and numbers are on the board. If not, close scope in one minute: "I'll lock these and assume the rest."
- Minute 20: the diagram is complete end to end. A box you haven't detailed stays a labeled box. A complete skeleton with one black box beats a detailed half.
- Minute 35: at least one deep dive is finished. Close it with its main trade-off, even if there is more to say.
For a 60-minute slot the marks are 13, 25 and 48.
Behind schedule? Compress, don't skip. Two minutes of scope beats no scope, and a one-minute wrap-up beats none.
The interviewer owns the agenda and may change it: "skip the estimates," "let's go straight to the storage layer." Follow them. The budget is your default when nobody says otherwise, not a plan to defend.
Make the plan audible
Glance at the clock at each transition and say where you are: "That's the flow end to end. We're about 20 minutes in, so I'd like to go deep on location ingestion now." A transition sentence costs five seconds and keeps the interviewer oriented. Keep it plain. "I will now enter phase three" sounds like a script, and the interviewer hears you reciting instead of designing.
Your turn: it is minute 18 of 45. You have eleven requirements on the board and the interviewer has not redirected you. What do you say, and what happens to the rest of your time?
Check your answer
Pick the three or four functional requirements and the one or two qualities that drive the design, say so, and move: "Let me lock these five. The rest I'll treat as out of scope unless you want one of them. I'm going to sketch the end-to-end design now." Park the other six on an "out of scope, maybe later" line.
You have 22 minutes before your questions and you are 8 minutes behind. Take most of that from the sketch and the wrap-up, and protect the deep dive: sketch from 18 to 25 at box level, deep dive from 25 to 38, wrap-up from 38 to 40. Do not ask the interviewer to choose the five for you. Prioritizing is part of what is being scored.
Change one requirement: at minute 2 the interviewer says, "We're short on time today. We have 35 minutes, including your questions." Re-plan.
Check your answer
Keep the shape and shrink everything a little: intro 0β2, scope 2β7, sketch 7β15, deep dive 15β27 (one topic), wrap-up 27β30, your questions 30β35. That is 2 + 5 + 8 + 12 + 3 + 5 = 35. Your checkpoints become 7, 15 and 27, and the deep dive still gets the biggest share. Say the new plan in one sentence so the interviewer knows you heard them.
Move 4: Breadth first, then depth where the risk is
Smell: you have spent twelve minutes on the schema, there is no API, and nobody knows how a request gets from the phone to the database.
The skeleton test
Before any box gets detail, every box on the path of a request should exist. Test that by tracing one write and one read out loud. Here is the ride-hailing sketch after scoping:
driver app βββ WebSocket βββΊ driver gateway ββ positions every 4 s βββΊ location service
β² β
β offer to driver βΌ
β accept back latest positions in memory,
β one shard per city
βΌ β²
rider app ββ "find me a ride" βββΊ matching service ββ nearby-driver query ββββββ
β² β
ββββ "driver on the way" ββββββββββββββ€
βΌ
trips DB (Postgres): trips, payments, profiles
Two things are left out to keep it readable: the load balancers in front of each service, and the append-only log that keeps on-trip pings as route history.
The write: a driver's app sends a position over its WebSocket, and the gateway passes it to the location service, which overwrites that driver's entry in the shard for their city. The read: a rider asks for a ride, the matching service asks the rider's city shard for nearby drivers and offers the trip to one of them through the gateway, and when the driver accepts, the trip is written to Postgres and the rider hears "driver on the way." One sentence each, and the interviewer knows the whole system exists. Now depth has something to attach to.
The API and data model belong in this phase too, at the level of "the three endpoints and the two tables that matter." Design Process Steps walks through that step by step.
β οΈ Over-indexing on one area is the failure this test prevents. A candidate who loves databases can spend 25 minutes on a beautiful schema and never show how requests arrive, how clients get answers, or what happens when a node dies. The interviewer leaves without knowing whether they can reason about the whole system.
Choosing the deep dive
Pick it by risk, not comfort. Ask yourself:
- Which stated requirement is hardest to meet?
- Which number is largest?
- Where would a bug cost money or trust?
- What does this prompt have that every other design doesn't?
| Prompt | Deep dive worth taking | Why that one |
|---|---|---|
| Ride-hailing | Location ingestion and the nearby-driver search | 125,000 writes per second of data that is stale in 4 seconds |
| Concert ticketing | Seat holds during the on-sale rush | "Never sell a seat twice" meets the biggest spike |
| Metrics and monitoring | The write path, downsampling and retention | Huge write volume; reads are time ranges |
| Collaborative document editing | Merging concurrent edits | Two people typing in the same paragraph must end up with the same text |
The workload shape points you somewhere too. Read-heavy systems usually need depth on caching, replicas, CDNs and invalidation. Write-heavy systems need depth on partitioning, batching, logs and queues. And watch for write-critical systems, such as payments and bookings. The volume may be tiny, but correctness under retries and concurrency is the whole deep dive.
Then say why you chose it, and check once:
"I'd like to go deep on location ingestion. It's 125,000 writes a second and matching depends on it. Unless you'd rather see matching?"
If the interviewer picks something else, go there. At senior levels you are expected to propose the deep dive; the interviewer can still overrule you.
Predict first: a chain of 2,000 clinics wants a reminder sent 24 hours before every appointment. There are 300,000 appointments a day. A reminder may be a few minutes late, but it must never be skipped, and a patient should not get it twice. Where is the risk?
Check your answer
Not in the volume: 300,000 a day is about 3.5 reminders per second. The risk is reliability: finding every reminder that is due without scanning the whole table each minute, sending each one at least once when the SMS provider fails or a worker crashes, and not sending it twice. So the deep dive is the scheduling and dispatch path: due-time indexing, retries, and a deduplication key per appointment. Name that key's limit too: a crash between the SMS send and recording the key can still send a rare duplicate, unless the SMS provider accepts an idempotency key. The clinic admin screens are not the deep dive. The biggest risk is not always the biggest number.
Gold-plating: complexity needs a requirement
Here is a prompt: "Design an employee directory for a company of 5,000 people." One candidate proposes eight microservices, Kafka for profile updates, Redis Cluster, and multi-region replication with automatic failover.
Run the numbers. 5,000 people Γ 20 lookups a day is 100,000 lookups a day. Squeeze them into an 8-hour workday and that is about 3.5 per second. Triple it for the morning rush and it is about 10 per second. The whole dataset, at 5 KB a person, is 25 MB. At that load, every extra component is pure cost: something more to deploy, monitor, secure and get paged about at 3 a.m. One service and one relational database with a replica does the job, and the database's own text search can handle names.
Every component has to pass one test: "I'm adding X because of requirement Y, and it costs Z." If you cannot finish that sentence, X goes. Extreme Programming has a name for the same instinct: YAGNI, "You Aren't Gonna Need It."
Scale thinking still belongs in the interview, as an evolution path, not as the starting design: "At 10Γ the lookups this still fits one database. If the company were a hundred times bigger, I'd add a cache in front of profile reads first." That shows you know what changes and when, without paying for it now.
Move 5: Read the interruption, answer in the trade-off shape
Smell: the interviewer asks a question, and your first thought is "I've been caught."
Most questions are not corrections. They come in three kinds:
| Kind | Sounds like | What they want | Your move |
|---|---|---|---|
| Steer | "Nice. What happens to a trip if the rider's phone loses signal?" | Coverage of something they need to score | Answer it, then continue from there |
| Probe | "You said in-memory. How exactly does the nearby-driver query work?" | Proof you understand the box you drew | Go one level deeper, with specifics |
| Challenge | "Why not keep positions in Postgres?" | Whether your choice has a reason, and whether you can update it | Give the requirement and number behind it, or change your mind and say why |
If you can't tell which kind it is, ask: "Do you want me to go deeper on this now, or note it and carry on?" That is not weakness. You are agreeing on priorities with the person who sets them.
The trade-off shape
Every decision you state out loud has five parts: options β the requirement that decides β choice β cost β check. For example, choosing how ride events reach two other teams:
"Ride events need to reach billing and analytics, and both need every event. I see two options. A queue broker can do that with one queue per consumer (a fanout exchange, for example), deleting each message once it's acknowledged. A log like Kafka lets each consumer group read every event at its own pace and re-read old ones. What decides it is replay: after a billing bug we'll want to re-read last week's events, and only the log keeps them, so I'd pick the log. The cost is one more system to run, plus sizing retention and managing partitions and consumer groups. Either way, consumers must tolerate duplicates, because a crash between processing and acknowledging redelivers. Does that fit what you had in mind?"
Likely follow-up: "What if analytics falls a day behind?" Have the answer ready: a log keeps events for a configured retention period (seven days by default in Kafka), so analytics catches up on its own without billing noticing. Size the retention to cover the longest outage you are willing to tolerate.
The short form is two sentences: "I'll use X because of requirement Y. It costs Z, which is acceptable becauseβ¦" Compare that with narrated anxiety: "Storageβ¦ there are lots of optionsβ¦ SQL, NoSQLβ¦ it dependsβ¦" Both admit uncertainty, but only the first one moves the design forward.
A technology name with no reason attached is a magic word. "We'll add Kafka" tells the interviewer nothing except that you have heard of Kafka.
On silence: "Let me think about that for twenty seconds," followed by twenty seconds of thought, is fine. It usually produces a better answer than an instant ramble. Two minutes of unexplained silence is different. The interviewer cannot tell thinking from being stuck, so they have nothing to score.
When a clean answer is wrong
The interviewer says: "You shard users by user ID. A pop star with 80 million followers joins, and their profile gets hammered. What now?" A candidate answers:
"That's the hot-shard problem. I'd use consistent hashing with virtual nodes to spread the load more evenly across shards. If monitoring later shows specific hotspots, I'd move those users to dedicated shards."
It has the right shape: it names the problem, offers options and makes a recommendation.
Predict first: what is wrong with it?
Check your answer
The recommended first fix does nothing for this problem. Consistent hashing with virtual nodes spreads many keys evenly across nodes, and it reduces how much data moves when you add a node. But every request for one user ID hashes to the same point on the ring, so it still lands on one shard. The pop star's traffic stays exactly where it was.
A fix aimed at the hot key looks like this. The traffic is almost all reads, so serve the pop star's profile and recent posts from a cache in front of the shard; read replicas of that shard help too. If one key's writes were the problem, such as a like counter, split it into N sub-keys and add them up on read. Detect hot keys from per-key traffic metrics instead of keeping a list of celebrities by hand. The costs: a cache can serve briefly stale data, and split keys make every read touch N places.
The lesson: the answer had a perfect shape and a false center. Good structure is what made the error easy to spot. Hot keys are covered in depth in Databases & Storage.
Driving without steamrolling
π§ You hold the pen; the interviewer holds the priorities. You propose the plan, the next step and the deep dive. They can redirect you at any moment, and when they do, you follow.
Check in at natural seams: after scope ("does this match what you had in mind?"), after the sketch ("I'd go deep on X next, or is there something you'd rather see?"), and before a long detour.
Two failure modes sit on either side of that:
- Passenger: asking "what should I do now?" at every step. The interviewer has to drive, and that is exactly what they were scoring you on.
- Monologue: fifteen minutes without a pause. When the interviewer finally cuts in, you have committed to decisions they wanted to explore differently, and you missed every hint they tried to give.
Listen for two kinds of signal. Depth signals mean go deeper: a follow-up on what you just said, your own term repeated back ("you said eventual consistencyβ¦"), "say more." Redirect signals mean move on: "sounds good, what aboutβ¦", "assume that's handled", a new question before you finish. Close your point in one sentence ("there's more to say on caching; I'll come back to it if we have time") and follow.
Tangents: whose idea was it?
This one question tells you what to do with a deep side topic.
- The interviewer raised it. Answer at the depth they asked for, in a minute or two, then offer the choice: "I can go deeper on how drivers' apps reconnect now, or finish the trip flow first. Which is more useful?" It is their priority. You are just making the time cost visible.
- You raised it. You noticed something fascinating on the way. Write it on a "later" list on the board and keep going. Come back to it in the deep dive or the wrap-up.
Your turn: classify each question from the interviewer, then answer the third one in two sentences.
- "How does the location service know which city shard to write to?"
- "Let's set matching aside. What happens to a trip in progress if the rider's phone loses signal?"
- "Couldn't one Postgres table with a spatial index hold the positions?"
Check your answer
1 is a probe (prove you understand the box you drew). 2 is a steer (they want coverage of a failure case). 3 is a challenge.
A two-sentence answer to 3: "At 125,000 updates a second, each also rewriting an index entry, I don't want positions on the same primary as trips and payments, and for matching a position is worthless four seconds later, so memory without durability fits better. If we had a thousand updates a second, I'd happily keep them in Postgres." The second sentence is what makes it strong. It names the condition under which you would choose differently, which shows the choice came from the numbers and not from habit.
Move 6: Stuck, wrong or out of time: recover out loud
Smell: you have been silent for a minute, or you just saw a hole in your own design.
You don't know a fact
Say what you would assume, why, and whether the design depends on it:
"I don't remember [X] exactly. I'd expect it to be around [Y], because [reason]. The design only changes if it's below [Z], so I'll assume Y and flag it as something to verify."
The opposite is bluffing, and a bluff falls apart on the next probe. A candidate says, "Postgres can't do more than 1,000 writes a second, so we need Cassandra." The interviewer asks where the 1,000 comes from, and the biggest decision in the design goes with it. A single Postgres primary on current hardware handles far more than that for small writes. The honest version keeps the design and the credibility: "I don't know the exact ceiling on our hardware. I'd expect one primary to take our peak of 3,000 small writes a second. If a load test says otherwise, partitioning by customer is my next step."
You don't know what to do next
Three questions always have an answer:
- Which stated requirement does my design not yet clearly meet?
- What happens when each box dies? Walk the diagram from left to right.
- What breaks first at 10Γ the load?
These are also the cure for running out of things to say at minute 20. If you finish the design early, the remaining minutes belong to these three questions.
You find your own mistake
Say it the moment you see it, with the fix and its effect:
"Wait, I want to fix something. My nearby-driver query reads every driver in the city and computes the distance to each, and the same query feeds the 'cars near you' map. Say 10,000 riders in a big city have the app open, each refreshing every 5 seconds: that's 2,000 searches a second. With 50,000 drivers online in that city, that's 100 million distance checks a second. I'll index positions by geohash cell instead and check only the rider's cell and its eight neighbors. At about 50 drivers per cell, that's 450 checks per search, or 900,000 a second, which is cheap. It changes only the location store; the rest of the diagram stands."
Self-correction is a good signal. It is exactly what you would want from a colleague in a design review. Hiding the mistake and hoping is the worst option, because the interviewer has probably seen it already.
The interviewer disagrees
Find out what they are pressing on: "Is your concern the write load or the operational cost?" Then either defend your choice with the requirement and the number behind it, or change the design and say what changed your mind. Both are good outcomes. Digging in without a reason is not. Mock Interviews & Communication practises handling pushback in depth.
You are out of time
Don't rush a half-finished deep dive into something incoherent. With two minutes left, say: "Let me stop there and summarize: the requirements, how the design meets each one, the biggest risk, and what I'd do next." Sixty seconds of summary lets the interviewer score the whole design instead of whatever sentence you were in when time ran out.
Your turn: it is minute 31 of 45, and you are in the deep dive on matching. To cut wait times, you have just changed the design to offer each request to the three nearest drivers at once. Then you realize two drivers can accept the same trip. Write what you say, in three or four sentences.
Check your answer
Something like: "I've spotted a bug: if two drivers accept the same request, both would think they got it. I'll make acceptance a conditional write, where the trip's driver is set only if it is still empty, so the second driver gets 'trip already taken.' That's one atomic update on the trip row, it costs the second driver a wasted tap, and it doesn't change anything else, so I'll carry on with the deep dive."
The three parts are: name it plainly, give a fix that is actually atomic (a check followed by a separate write would have the same race), and say what it costs and how it affects your time.
Final round: no label on the situation
Real interviews don't announce which move they need. For each situation, decide what you do before you open the answer.
Situation 1: minute 0
"Design a service that lets a bank's 3 million customers download their monthly statements as PDFs." What are your first ninety seconds?
Check your answer
Restate the prompt in one sentence, then ask the questions that change the drawing. Are statements generated in advance or on demand? When do downloads peak? How long must statements be kept? Who may see a statement, and how is that checked? State your assumptions for anything the interviewer leaves to you. Suppose statements are generated in advance on the 1st, within 6 hours, at about 200 KB each: that is 3,000,000 Γ· 21,600 s β 140 PDFs a second, and 600 GB a month, or 7.2 TB a year. So a steady batch job writes the PDFs to object storage, which is where 7 TB a year belongs, and a download is a permission check plus a read from storage. Then announce the plan: sketch, then a deep dive on access control, because a statement shown to the wrong customer is the failure that matters most. The generation pipeline is a routine batch job at 140 PDFs a second; in a 60-minute slot it could be the second deep-dive topic.
Situation 2: minute 16 of 45
You are in the middle of the sketch. The interviewer frowns and asks, "What if two of these requests arrive at the same time?"
Check your answer
It is a probe with a hint in it: they suspect a race. Don't brush it off until the deep dive. If you're unsure which requests they mean, ask. Then name the race in one sentence, give the fix at one level of depth, and offer to go deeper now. If this race sits on the core risk of the prompt, it is probably your deep dive anyway, so say so.
Situation 3: minute 26 of 45
You are deep in your cache design. The interviewer's last three questions were all about writes.
Check your answer
That is a redirect signal. Close the cache in one sentence with its main trade-off, say "It sounds like the write path is the more interesting part. Let me go there," and move. You have nine minutes until the 35-minute checkpoint.
Situation 4: minute 38 of 45
"Say traffic grows 10Γ. What breaks first?"
Check your answer
Go to your numbers, not your gut. Multiply each box's load by ten, compare it with what that box can take, and name the first one that crosses its limit, with the move and its cost, in a few sentences. For the ride-hailing design: "Position writes go from 125,000 to 1.25 million a second, and a big city's single shard would have to take ten times its current writes. That's the first thing I'd expect to break, so I'd split big cities into several shards by geohash region instead of one shard per city. The cost is that a search near a region border has to ask more than one shard. Trips average about 116 a second, so about 350 at a 3Γ peak; at 10Γ that is about 3,500 new trips a second, and at five writes per trip about 17,000 writes a second. That's more than I'd trust to one primary, so the trips database is the second thing I'd plan for: partition trips by city."
Situation 5: minute 4 of 45
The interviewer says, "Let's skip requirements. Assume whatever's reasonable."
Check your answer
Take the gift, but spend thirty seconds saying your assumptions out loud: three core features, two numbers and the one quality that matters most. Without them the design has nothing to be judged against, and neither of you can tell whether it is right. Then go straight to the sketch.
How to practise
Knowing the moves is not the same as doing them at minute 26 with someone watching. Practise the moves the same way you'll use them:
- Run timed, spoken sessions. Set a 45-minute timer, pick a prompt you have not designed before, and talk the whole time as if the interviewer were there. Record the audio. Silent practice in a notebook trains the wrong habit, because communication is one of the scored areas.
- Check the checkpoints. Were you at 10, 20 and 35 on time?
- Keep a short log after each run, and choose the next session's focus from it:
Prompt: parking-garage reservations Slot: 45 min
Scope 3-11 (8) target 3-10
Sketch 11-22 (11) target 10-20
Deep dive 22-36 (14) target 20-35 topic: holding a spot without overbooking
Wrap-up 36-40 (4) target 35-40
Silent > 30 s: once, while choosing how to expire unpaid holds
Requirement never shown to be met: "garage owners see bookings in real time"
Next time: trace the garage-owner read path during the sketch
- Mock with a partner who interrupts. Ask them to steer, probe and challenge at random, and to redirect you mid-sentence at least once.
- Practise on the real tool. In a video interview you will probably draw in a shared whiteboard app, which is slower than a pen. Learn its shortcuts before the day.
The framework will feel mechanical for the first few runs. That is expected. After a while you stop thinking "which phase am I in?" and your attention goes to the design.
Cheat sheet: pressure β move β cost
| When you notice⦠| Move | What overdoing it costs |
|---|---|---|
| You are drawing before anyone has said a number | Scope: questions that change the design; write down assumptions | Fifteen minutes of questions leave no deep dive |
| A number with no "soβ¦" | Estimate only what decides something | Skipping the one number that decides hides the hard part |
| Minute 10, 20 or 35 goes by | Check the clock; compress, don't skip | Robotic phase announcements |
| You are perfecting one box | Skeleton first, traced end to end; deep dive chosen by risk | A skeleton with no depth anywhere |
| A component you can't justify | "X because of Y, costs Z", or drop it | Under-building what a stated requirement needs |
| An interruption | Steer, probe or challenge? Answer in the trade-off shape | Asking permission for every step |
| Silence, a hole in the design, or a disagreement | Recover out loud: assume, name, fix, summarize | Over-apologizing instead of fixing |
Before moving on, pick a prompt you have never designed and explain aloud, in two minutes: how you would spend a 45-minute slot and where your three checkpoints are, which deep dive you would choose and why, how you would state your biggest decision in the trade-off shape, and what you would do if you went blank at minute 25. Then run one full timed session and log it.
Next: Requirements Gathering turns Move 1 and Move 2 into a repeatable conversation. Design Process Steps covers the sketch and the deep dive: API, data model, diagram and bottlenecks. When you want an end-to-end playbook for a prompt you have never seen, go to Realistic Design Examples. For timed mocks and pushback, see Mock Interviews & Communication.