Mock Interviews & Communication

Run a timed mock end to end on a fresh prompt, state trade-offs clearly, handle pushback and changing requirements, and score yourself honestly

Last generated

Lesson 18 of 18 available15 practice questions

SPACED REPETITION · 15 practice questions

Make this lesson stick.

Try 3 questions now. No account needed. Sample answers aren't saved.

You know the design. Can you deliver it in 45 minutes?

Picture a candidate who has done all the reading. She can explain consistent hashing, leader election and Kafka consumer groups. The prompt is: "Design a service that runs scheduled jobs for other teams: cron, but for the whole company." Here are her 45 minutes, as the interviewer's notes would record them:

Clock What happened
0:00–0:13 Twenty-one clarifying questions, including the UI, log retention, alert emails and billing. No numbers written down.
0:13–0:27 Eleven boxes drawn in near silence: gateway, auth, Kafka, ZooKeeper, Cassandra, Kubernetes, Elasticsearch…
0:27 Interviewer: "What happens at 9:00, when 40% of the hour's runs are due at once?" Candidate: "We scale horizontally."
0:34 Interviewer: "Why Cassandra for the schedule?" She apologizes, erases Cassandra and writes "Redis?"
0:36 She starts on Elasticsearch for job logs.
0:40 The interviewer moves on to her questions. There is no path from "a job is due" to "a run has started", and she gave one reason in 40 minutes of design.

Nothing she said was false. The interviewer still can't recommend a hire, because the evidence never appeared: a design that works end to end, a reason behind each choice, and a steady response to a challenge. An interviewer can only score what they saw and heard.

Knowing a design and delivering it are different skills, and the second one improves only by doing it: out loud, against a clock, with someone pushing back. That's a mock interview. This lesson shows you how to run one on a fresh prompt and what to practice inside it.

Look at what broke, in order: the clock (the design never got finished), the reasons (choices arrived without trade-offs) and composure (pushback turned into retreat). Each has a move, and three more cover the rest of the session:

# Pressure in the room Move What it sounds like
1 Minute 25 and no end-to-end design Run the clock "It's 0:20 and a write and a read are traced end to end; now the deep dive."
2 "I'd use Redis" and nothing else Say the trade-off in four parts "X, because…; it costs…; I'd switch if…"
3 Eleven boxes nobody can follow Make the board do the remembering "Here's the run path; the numbers are on the left."
4 "Why not X?", a hint, a failure question Read the interjection, then keep or change "Let me check that against the 2,000 writes a second."
5 A requirement changes at minute 30 Re-derive out loud "That changes the claim, not the schedule. New number: …"
6 The mock ends with only a feeling Score it with evidence "At 0:12 I added queues without a reason."

This lesson runs on the course's single interview framework from Interview Framework & Strategy (Scope, Sketch, Deep dive, Wrap-up) and assumes the questions from Requirements Gathering and the steps from Design Process Steps. Here you rehearse the whole thing under pressure.

What turns practice into a mock

Five ingredients. Drop any of them and you're studying: useful, but a different skill.

  • Out loud, the whole time. Thinking silently doesn't train talking.
  • A clock you don't pause. Stopping to look something up ends the rehearsal.
  • Someone who knows things you don't: a prompt card with hidden facts and planned pushback.
  • A record: audio, video or timestamped notes, so the debrief runs on evidence.
  • A debrief that ends in one change for the next mock.

Move 1: Run the clock

Pressure: it's minute 25, and there's still no path on the board from request to result.

Your mock runs on the course's clock, the one Interview Framework & Strategy defines. Don't invent your own split; rehearse this one until it's automatic:

Phase What you produce 45-minute slot 60-minute slot
Intro and prompt Nothing yet 0–3 0–5
1. Scope Features, qualities, two or three numbers, assumptions 3–10 5–13
2. Sketch API, data model, and a diagram with one write and one read traced end to end 10–20 13–25
3. Deep dive The riskiest part designed properly, with its failure modes (one topic in 45 minutes, two in 60) 20–35 25–48
4. Wrap-up Summary, top risks, what changes at 10× 35–40 48–55
Your questions Your questions for the interviewer 40–45 55–60

In a 45-minute mock you watch three checkpoints: 0:10, scope closed; 0:20, the sketch complete end to end; 0:35, at least one deep dive finished. The one that matters most is 0:20. A deep dive on a design that doesn't work end to end is polishing one part of a machine that can't run yet, and if time runs out mid-dive, you still leave a working design behind.

Three rules keep the clock honest:

  • Behind schedule? Compress, don't skip. A shorter Scope with assumptions said aloud beats a missing sketch; a shorter deep dive beats a missing wrap-up.
  • The interviewer outranks the plan. If they steer you somewhere, go (Move 4). The checkpoints are for when nobody is steering.
  • A longer slot buys depth, a second deep-dive topic, not a much longer Scope (it grows by one minute).

Predict first: it's 0:14. You've asked your questions but computed nothing, and nothing is drawn. The interviewer is quiet. What do you do in the next 60 seconds?

Check your answer

Say the time and your plan out loud: "I'm behind, so I'll assume schedules are in UTC and a run may start up to a minute late, compute the one number that decides the architecture, how many runs are due at the top of the hour, and draw the path." Then do exactly that. That's compress, don't skip: Scope shrinks to assumptions and one number, but it still happens. Stating assumptions instead of asking is how you buy time back, and saying them aloud lets the interviewer correct one in a sentence. A long apology or a silent rush both burn the minute you need.

The prompt card

A mock needs someone who knows more than you do. That person holds a prompt card: the prompt, the facts they reveal only when you ask, the pushback they'll deliver, and one mid-session change. Here is the card for this lesson's running example:

PROMPT (read aloud)
  "Design a service that runs scheduled jobs for other teams:
   cron, but for the whole company."

HIDDEN FACTS (reveal only when asked)
  - 2 million runs a day. About 40% of each hour's runs are due at the
    top of the hour (hh:00:00); the rest are spread out.
  - A run must start within 60 seconds of its scheduled time.
  - Runs last about 3 minutes on average; some last hours.
  - The service only triggers runs; each team's own workers execute them.
  - A failed run may be retried, but two runs of the same job must never
    overlap, and a due run must never be silently skipped.
  - Out of scope unless asked: the UI, deploying job code, job output.

PLANNED PUSHBACK
  - During the deep dive: "Why not keep the schedule in a Redis sorted set?
    It's faster."
  - "A worker freezes for 90 seconds in a garbage-collection pause while it
    holds a run. Then what?"
  - "Isn't your database a single point of failure for every team's jobs?"

MID-SESSION CHANGE (around 0:30)
  - "Each team may have at most 200 runs going at once."

If you're the interviewer: reveal facts only when asked, and let a reasonable assumption stand. Deliver the planned pushback even when the design is right, because the candidate is practicing the response. Give one hint if they're stuck for more than about two minutes. Write timestamps. Don't teach during the mock; save it for the debrief.

Practicing alone? Have a friend write cards, or write a batch yourself and use each one weeks later, when you've forgotten the details. Keep a visible timer, record audio, and read the pushback lines from the card at the planned points.

The run, traced

Here is a strong run on that card: what was said or drawn, when, and why it counts.

Clock Said or drawn Why it counts
0:02 "Company-wide cron: the hard parts are probably the top of the hour and never running one job twice at once. Plan: scope until 0:10, a sketch by 0:20, one deep dive on the riskiest part, wrap-up at 0:35." The plan is public, so the interviewer can redirect early
0:03 Scope: how many runs a day, and how bunched? How late may a run start? What if the previous run is still going? Who runs the code? Every answer changes a decision
0:07 Numbers box: 33,333 runs due at hh:00:00, 556 claims/s to start them within 60 s, about 2,000 writes/s at the worst moment Shows that the average is a trap and one database is enough
0:10 Sketch: API and two tables, jobs and runs Scope closed on time
0:12 Diagram, narrated one path at a time: triggering a run (a write) and a team's run history (a read); a queue per team appears Sketch complete end to end at 0:19
0:20 "For the deep dive I'd take the run claim, because a job running twice at once is the failure teams can't tolerate. The other candidate is the top-of-hour burst. Which would you prefer?" Interviewer: "The claim." Chosen where the requirements bite, with the interviewer's say
0:21 The claim as one conditional write with a lease; "why not Redis?" answered with numbers; "why the queues?" at 0:24 A mechanism plus a defended trade-off
0:26 Failure modes: a worker frozen past its lease (a fence number); the primary dying (a synchronous standby) A real gap, fixed out loud
0:30 Change: at most 200 running runs per team. Re-derives: a counter row, and 200 is below the largest team's average load (about 420 runs at once) Updates without losing the thread
0:35 Wrap-up: three decisions and their costs; top risks, including "isn't the database a single point of failure?" (fail closed); what changes at 10× Deep dive finished on time; a clean close
0:40 The candidate's questions for the interviewer The design ended on time, so these minutes survive

The numbers at 0:07

Scope ends with two or three numbers that decide something. These come from the card plus two design assumptions said aloud: the scheduler records each trigger with one write, and a running job renews a lease every 30 seconds.

On the board Arithmetic Decision it drives
33,333 runs due at hh:00:00 2M runs ÷ 24 hours × 40% Design for the top of the hour, not the average
556 claims/s to start them within 60 s 33,333 ÷ 60 s The claim path's rate; any queue carries it
About 2,000 writes/s at the worst moment In the first minute, 556 claims + 556 trigger records a second; then up to 1,111 renewals (33,333 ÷ 30 s) + about 185 finishes (33,333 ÷ 180 s) One relational primary is plenty

The average, 23 runs a second (2 million ÷ 86,400), is 24 times smaller than the top-of-hour rate, which is exactly why it's a trap. The writes don't all land in the same second: the claims and trigger records come in the first minute, and the renewals peak just after it, while the runs thin out as they finish. Added up second by second (with 3-minute average runs, plus the runs spread over the rest of the hour), the peak comes to about 2,000 a second; adding every term at its maximum gives about 2,400. Both fit one primary.

💡 The sentence that changes the interview: "The average, 23 runs a second, is a trap. The design is sized by one minute an hour: 33,333 runs due in the same second, 556 claims a second, and about 2,000 database writes a second at the worst moment. That fits one PostgreSQL primary, so no sharding." Say it at 0:07 and nobody spends ten minutes partitioning a jobs table that fits on one machine.

Notice what is not in the numbers box: storage. A run record of about 500 bytes, 2 million times a day, is 1 GB a day; thirty days of history is 30 GB. Computing it would take two minutes and decide nothing. Compute the numbers that change a decision.

Where the clock goes wrong

  • Requirements sprawl. Twenty questions, including ones whose answers change nothing ("What color are failed runs in the UI?"). Ask the ones that change a decision; assume the rest out loud.
  • A deep dive before the design works. At 0:08 you're explaining how to parse cron expressions. Write "cron parsing" on the parking list (Move 3) and keep drawing.
  • No wrap-up. Time ends mid-sentence, and the last thing the interviewer heard is a detail. At 0:35 you start the wrap-up, whatever is happening.
  • Refusing to be steered. If the interviewer takes the wheel at 0:15, the plan bends. Follow them, and pick the checkpoints up again when they let go.

Your turn: the interviewer says this mock will run 60 minutes. Using the 60-minute column, what does the extra time buy on this card, and where are your checkpoints?

Check your answer

It buys a second deep-dive topic; Scope grows by only a minute and still closes by 0:13 and the sketch by 0:25. The deep dive runs from 0:25 to 0:48: the run claim first, then the top-of-hour burst (several scheduler instances sharing the due jobs, and enqueueing a run before recording it as triggered). The wrap-up runs from 0:48 to 0:55, and the last five minutes are your questions. If you let Scope stretch into the extra time, you get a more thoroughly scoped problem and the same single, rushed deep dive.

Move 2: Say the trade-off in four parts

Pressure: you say "I'd use Postgres here" and move on. Or you say "it depends" and stop.

Naming a component is not evidence of judgment; anyone can name a database. The evidence is the reasoning, and it has four parts:

Decision: what you chose. Because: the stated requirement or number it serves. Cost: what you give up, with a number if you have one. Switch if: the condition that would make you choose differently.

As a sentence: "I'll use X, because [requirement or number]. It costs [what we give up]. If [condition], I'd switch to Y." It takes about twenty seconds to say, and it answers the interviewer's next three questions before they ask them.

Three upgrades from the mock

Leases.

  • Weak: "Workers send heartbeats."
  • Strong: "A worker's lease on a run lasts 90 seconds and it renews every 30, because with up to 33,333 runs in flight after the top of the hour that's 1,111 renewals a second, which one primary takes. It costs up to 90 seconds before a dead worker's run is retried. If teams needed faster recovery, I'd use a 30-second lease renewed every 10 seconds, at up to 3,333 writes a second."

Replication.

  • Weak: "Postgres with a replica for high availability."
  • Strong: "Jobs and leases live on one Postgres primary with synchronous replication to either of two standbys, because a claim we've acknowledged must survive the primary dying: lose it, and a second worker can start the same job. It costs a round trip to a standby on every commit, typically low single-digit milliseconds inside one region and nothing at about 2,000 writes a second, and if both standbys are down, commits wait until one returns. For run logs, where losing the last second is harmless, I'd replicate asynchronously and skip the round trip."

In PostgreSQL that is the quorum setting synchronous_standby_names = 'ANY 1 (a, b)': each commit waits until at least one of the two standbys has written the commit record to durable storage. The PostgreSQL docs are blunt about the cost: the minimum extra wait is the round trip between primary and standby, and commits may never complete if too few synchronous standbys are available.

"It depends."

  • Weak: "What happens when a run is still going at the next due time depends on the job."
  • Strong: "It depends on one thing: whether a missed run matters. For 'refresh a cache', skipping one is fine; for 'send the hourly invoice batch', it isn't. By default I'll skip the new run, record it and notify the owning team, because most periodic jobs catch up on their next run, and a skip nobody hears about is the silent skip the requirements forbid. It costs noise: a job that keeps overrunning sends a notice every hour until someone fixes it. Jobs that can't miss a run can opt into 'queue one behind'."

"It depends" is the start of a good answer, not the answer. Name the variable it depends on, pick a default, and say what would flip it.

The mock's decisions, in four parts

Decision Because Cost Switch if
A queue per team between the scheduler and the workers Teams' workers take runs at their own pace; 556 runs/s at the top of the hour Duplicate deliveries; one more system to run Teams want to be called: push over HTTP with retries
Job state and leases in Postgres, claimed with a conditional write "Never overlap", "never skip"; about 2,000 writes/s A primary to operate; every claim goes to it Writes far beyond one primary: partition
Lease 90 s, renewed every 30 s Up to 1,111 renewals/s at the worst moment Up to 90 s before a dead worker's run is retried Faster recovery needed: 30 s lease, up to 3,333 renewals/s
Synchronous standby (any 1 of 2) Acknowledged claims must survive a failover A commit round trip; writes wait if no standby is up Data that is harmless to lose: asynchronous
Overlap default: skip the new run, record it, notify the team Most periodic jobs catch up next time; no skip is silent A notice every time; a chronically slow job is noisy Jobs that can't miss a run: queue one behind

Predict first: what's missing from this? "I'll put Kafka between the scheduler and the workers, because it decouples them and it scales."

Check your answer

Almost everything. "Decouples" and "scales" don't point at a requirement or a number, and there's no cost and no switch condition. A four-part version: "Runs go through a queue per team, because teams run their own workers at their own pace and the scheduler must not wait for them; the queue also absorbs 556 runs a second at the top of the hour. It costs duplicates, because queues deliver at least once in the usual setup, so the claim must reject a run that has already started. At 556 messages a second any queue works, so I'd use whichever one we already operate rather than stand up a new Kafka cluster. If teams wanted us to call them instead, I'd push over HTTP with retries."

Traps

⚠️ Costs that aren't costs here. "Redis adds network latency", when a round trip inside one data center typically takes well under a millisecond and runs may start up to 60 seconds late. A cost only counts if it matters for these requirements.

⚠️ Costs that aren't true. "SQL doesn't scale horizontally" (sharded PostgreSQL and distributed SQL databases such as CockroachDB exist), or "NoSQL is eventually consistent" (DynamoDB offers strongly consistent reads). A wrong cost is worse than a missing one, because the interviewer will pull on it.

⚠️ Options without a choice. "We could use A, B or C" is a menu, not a design. Pick one, and put the runner-up in the switch condition.

⚠️ Hedging versus honesty. "Maybe we could possibly use a cache" sounds unsure of everything. "I don't remember Kafka's default retention; the property I need is a day of replay, so I'd set it explicitly" is precise about exactly one unknown.

What the interviewer asks next

Each part invites a follow-up, so have the number behind it ready:

  • After because: "What number makes that true?"
  • After cost: "How bad is that for a user?"
  • After switch if: "And what would you switch to?"

Your turn: the scheduler finds due runs by asking the database every second, instead of keeping an in-memory timer per job. Say it in four parts, with a number.

Check your answer

"The scheduler asks the database every second for jobs whose next run time has passed, using an index on that column, because that's one indexed range query a second, 86,400 small queries a day, and nothing is lost when a scheduler restarts: the database is the only state. It costs up to a second of delay plus the query, well inside the 60-second window; at the top of the hour one query returns 33,333 rows, so it pages through them. If runs had to start within milliseconds, I'd keep the next few minutes of due runs in memory with timers, rebuilt from the database on restart."

Move 3: Make the board do the remembering

Pressure: eleven boxes, arrows with no meaning, and the interviewer can't tell which path a due job takes.

The board is shared memory. If a number, requirement or decision isn't on it, both of you have to hold it in your heads while you talk about something else. Lay it out in three columns:

+-----------------------+--------------------------------+----------------+
| REQUIREMENTS          | DIAGRAM                        | LATER          |
| never overlap         |                                | cron parsing   |
| never silently skip   |   request paths, one action    | UI             |
| start within 60 s     |   at a time (below)            | multi-region   |
|                       |                                | job output?    |
| NUMBERS               |                                |                |
| 33,333 due at hh:00   |                                |                |
| 556 claims/s          |                                |                |
| about 2,000 writes/s  |                                |                |
+-----------------------+--------------------------------+----------------+

The left column is the contract. When you justify a choice, point at it: that's the because from Move 2. The right column is the parking lot. A tangent goes there in three words, so you can defer it without losing it and pick it up among the wrap-up's top risks.

Draw request paths, not a component catalogue

  1. One action at a time: register a job, trigger a run, run it, read its history. A path you can trace is a design; a pile of boxes is a list.
  2. Label every arrow with the operation and its rate. claim ~556/s at hh:00 tells the interviewer where the load is without another word.
  3. One sentence per box as you draw it: what it does and why it exists. If you can't say why, don't draw it. A Kafka box with no consumer is decoration.
  4. Keep the board current. When the design changes, change the drawing. A stale diagram contradicts you ten minutes later.
  5. Check in once, after the end-to-end path: propose your deep dive and say why.

Here is the board at 0:19, drawn as request paths. The bracketed notes are what the deep dive and the change add later, drawn in as they happen:

SCHEDULE PATH
  team --POST /jobs {cron, team}--> Jobs API --> Postgres: jobs (index on next_run_at)
  scheduler --every 1 s: due jobs, FOR UPDATE SKIP LOCKED--> Postgres
  scheduler --1. enqueue {job, run} (~556/s at hh:00)--> that team's queue
  scheduler --2. then advance next_run_at--> Postgres

RUN PATH (a team's worker)
  worker <--take {job, run}-- team's queue
  worker --claim: conditional UPDATE (~556/s at hh:00)--> Postgres
                                   [0:21: + 90 s lease, renewed every 30 s, up to ~1,111/s]
                                   [0:26: + fence number]
                                   [0:30: + team slot, in the same transaction]
  worker --finish--> Postgres      [0:26: must carry the current fence]

READ PATH
  team --GET /jobs/{id}/runs--> Jobs API --> Postgres: runs (last 30 days)

DATA
  Postgres primary                 [0:26: ==sync, ANY 1==> standby A | B]

Two details in that diagram carry requirements. The scheduler enqueues first and advances next_run_at second, so a crash between the two repeats a trigger instead of skipping one; the claim (below) rejects the repeat. And FOR UPDATE SKIP LOCKED lets several scheduler instances take due jobs without waiting on each other's rows; the PostgreSQL docs describe exactly this use, avoiding lock contention among consumers of a queue-like table.

In a remote interview with no drawing tool, type exactly this kind of path list into the shared editor, with the numbers at the top.

Pick the deep dive where the requirements bite

At 0:20 the Deep dive phase starts, with one topic in a 45-minute slot (two in 60). Pick the riskiest part, the one that:

  • carries the hardest stated requirement ("never overlap" → the run claim),
  • sees the biggest number (33,333 runs due in one second → the top-of-hour burst), or
  • the interviewer keeps pointing at.

Not the part you know best. If you love parsers, cron expressions are tempting, and they're the least interesting box on the board.

Predict: you get one deep dive. From these five (the job-definition API, the run claim, the top-of-hour burst, run-log storage, cron parsing), which do you propose, and how do you say it?

Check your answer

The run claim or the top-of-hour burst: one carries the hardest requirement, the other the biggest number. The strongest move proposes one with a reason, names the other, and lets the interviewer pick: "In the time we have, I'd go deep on the run claim, because a job running twice at once is the failure teams can't tolerate. The other candidate is the top-of-hour burst, 33,333 runs due in one second. Which would you prefer?" Whichever they don't pick goes into the wrap-up as a top risk. The API and run logs are routine, and cron parsing is a library call.

Deep dive: one WHERE clause keeps the promise

"Never overlap" comes down to one idea: make starting a run a conditional write on the job's row. The database gives the run a lease only if nothing else holds one (or the holder's lease has expired) and this run hasn't already finished, and it tells you whether it did. Every claim also bumps a fence number. Finishing works the same way: only the worker holding the current fence can record the run as done.

SQLite stands in for PostgreSQL here so you can run it (RETURNING needs SQLite 3.35 or newer; check sqlite3.sqlite_version). Run ids are full UTC timestamps, so they keep increasing across midnight; 'HH:MM' strings would make the next day's 00:00 run look older than yesterday's 23:00 and refuse it forever.

import sqlite3

db = sqlite3.connect(":memory:", isolation_level=None)   # autocommit
db.execute("""CREATE TABLE jobs (
    job_id        TEXT PRIMARY KEY,
    running_run   TEXT,               -- the run (a UTC timestamp) that holds the lease
    lease_until   REAL,
    fence         INTEGER NOT NULL,   -- goes up on every claim
    last_finished TEXT NOT NULL)""")
db.execute("INSERT INTO jobs VALUES ('hourly-report', NULL, NULL, 0, '2026-09-24T08:00Z')")

CLAIM = """UPDATE jobs SET running_run = :run, lease_until = :now + 90, fence = fence + 1
           WHERE job_id = :job
             AND (running_run IS NULL OR lease_until < :now)
             AND last_finished < :run
           RETURNING fence"""
FINISH = """UPDATE jobs SET running_run = NULL, last_finished = :run
            WHERE job_id = :job AND running_run = :run AND fence = :fence"""

def claim(job, run, now):
    row = db.execute(CLAIM, {"job": job, "run": run, "now": now}).fetchone()
    return row[0] if row else None          # the fence, or None if refused

def finish(job, run, fence):
    return db.execute(FINISH, {"job": job, "run": run, "fence": fence}).rowcount == 1

RUN = "2026-09-24T09:00Z"
t = 32_400.0                                  # 09:00:00, as seconds since midnight
a = claim("hourly-report", RUN, t)            # worker A
print(a)                                      # 1: A holds the lease
print(claim("hourly-report", RUN, t + 0.004)) # None: a duplicate message, 4 ms later
b = claim("hourly-report", RUN, t + 91)       # A froze; the lease expired; a retry takes over
print(b)                                      # 2
print(finish("hourly-report", RUN, a))        # False: A woke up holding a stale fence
print(finish("hourly-report", RUN, b))        # True
print(claim("hourly-report", RUN, t + 200))   # None: this run already finished

Why it holds under concurrency. In PostgreSQL's default Read Committed level, when two transactions update the same row, the second waits for the first to commit, then re-evaluates its WHERE clause against the new version of the row (PostgreSQL docs). If two workers take duplicate messages for the same run in the same millisecond, the second update waits, finds the lease taken, matches nothing and returns no fence. There's no "read, check in the app, then write" gap for a second worker to slip through.

Two production details, each worth one sentence in the interview:

  • Use the database's clock (now() in PostgreSQL) for lease_until, not each worker's. A worker whose clock runs fast would otherwise treat other workers' leases as expired early.
  • Expired leases need no cleanup to be safe: the WHERE clause already treats them as free. A background reaper only re-enqueues their runs so they get retried.

The sentence to say: "The invariant lives in two statements. Everything else, the scheduler, the queues and the workers included, can crash or repeat itself without breaking it."

Move 4: Read the interjection, then keep or change

Pressure: "Hmm. Why not keep the schedule in a Redis sorted set? It's faster."

Interviewers interrupt for different reasons, and they rarely say which:

They say Often means Your move
"What happens at 9:00, when 40% of the hour's runs are due?" A hint: the hard part is there Take it now, and say so
"Let's assume deploying job code just works." A scope cut One sentence, then move on
"Why not X instead?" A probe of your reasoning, or a hint that X fits better Compare X with your choice against the board; keep or switch
"Actually, we also need Y." New information Re-derive out loud (Move 5)
"What if Z fails?" A failure probe Trace it: what users see, the fix, its cost
"OK." or silence Often "move on", or they're writing Finish the sentence; check in if unsure

You can't reliably tell a probe from a hint by tone, and you don't need to. The move is the same: check the concern against the requirements and the numbers, then decide. Keep the design with reasons, or change it with reasons. What loses points is changing it without reasons (caving) or keeping it without reasons (digging in).

The response takes about a minute and has four steps:

  1. Restate the concern concretely: "You're asking whether Redis would make the schedule faster."
  2. Check it against the board.
  3. Decide: keep, modify or switch.
  4. Cost: what your decision gives up, and what would change your mind.

Pushback 1: a probe of a sound decision

Interviewer: "Why not keep the schedule in a Redis sorted set? It's faster."

You: "You're suggesting Redis for speed. Checking the board: we write about 2,000 times a second at the worst moment, which PostgreSQL handles, so speed isn't our constraint. The constraints are 'never overlap' and 'never skip', so an acknowledged claim, and every registered job, has to survive a failover. Redis replicates asynchronously, and its docs say even WAIT doesn't make it strongly consistent: acknowledged writes can still be lost in a failover. A lost lease means a second worker can start the same job. So the schedule and the leases stay in Postgres. The cost is a slower write path than Redis, which nobody notices at 2,000 a second. If the write rate grew by orders of magnitude, I'd partition the jobs across Postgres primaries before giving up durability. Redis could still hold a cache of the next minute's due runs, rebuilt from Postgres if it's lost."

The Redis claim comes straight from the Redis replication docs. Notice that the answer kept the design without dismissing the question: it said where Redis could fit.

Pushback 2: a real gap

Interviewer: "A worker freezes for 90 seconds in a garbage-collection pause while it holds a run. Then what?"

You: "That's a real gap. At 90 seconds the lease expires, another worker claims the run, and the frozen worker wakes up still running it. A lease tells a healthy worker when to stop: it times the lease on its own monotonic clock and stops well before expiry. A frozen worker can't check anything, though, and a pause can fall between a check and a write. The fence number is what stops its stale writes: our finish already refuses the old fence, and every write the job makes to shared storage must carry the fence, so the storage rejects any fence older than the newest it has seen. Jobs whose targets can't check a fence must be safe to run twice. Cost: teams pass the fence through or make jobs idempotent, and the two runs can still overlap in time; the fence only stops the stale writes from landing, and only where they're checked."

This is the fencing-token technique; Martin Kleppmann has described it with exactly this garbage-collection example. Here the right move was to change the design and say so plainly. Holding your ground on a real gap isn't confidence; it's a miss.

Pushback 3: partly right

This one arrives at 0:36, when the wrap-up lists the top risks:

Interviewer: "Isn't your database a single point of failure for every team's jobs?"

You: "Partly. The schedulers and workers are stateless. The state is one Postgres primary with a synchronous standby and automatic failover, so losing the primary costs a failover, not acknowledged claims, as long as failover promotes the most up-to-date standby (Patroni's quorum mode, for example, checks this). While no primary is available, I fail closed: no claims succeed, so nothing starts without a lease. Failing open, starting runs without one, would break 'never overlap' for every job at once. Cost: runs due during the outage start late, possibly past the 60-second window. Nothing is skipped, though, because next_run_at only advances after a run is enqueued, so the schedulers catch up afterwards. Is a late start during a failover acceptable to the teams? If not, the next step is a faster failover, not failing open."

Fail open or fail closed is a decision, not a detail, and the requirement picks it. Here, an unleased run could overlap with another, so closed wins. A rate limiter in front of a like button would fail open, because letting a few extra likes through hurts nobody.

Hints are gifts

If the interviewer asks about the top of the hour while you're polishing the job-definition API, drop the API: "Good, the burst is the harder problem. Let me go there now and come back to the API if there's time." Insisting on finishing your current thought costs you the hint and the time.

When you don't know

Say what you don't know, reason from what you do know, and say how you'd find out:

"I don't remember exactly when our queue redelivers a message that nobody has acknowledged. If it redelivers while a slow worker is still running the job, we get a duplicate delivery; the claim refuses it, because the lease is still held. So correctness doesn't depend on the answer, only wasted work does. I'd check the queue's docs."

That's one honest unknown, bounded and with a plan, instead of a bluff the interviewer can puncture.

Predict: classify each of these, and give your first sentence. (a) "We're at 35 minutes." (b) "What about daylight saving time?" (c) "Couldn't two workers start the same job?"

Check your answer

(a) A time signal: the deep dive's time is up. Close the current point in one sentence and start the wrap-up: "Let me stop the claim there and wrap up: the decisions, the top risks, and what changes at 10×."

(b) A probe, or a hint about schedules. One short question keeps you from answering the wrong version: "Are schedules in UTC, or in each team's local time?" In a zone with daylight saving time, one local hour doesn't exist in spring and one happens twice in autumn, so a local-time job needs a stated policy (run the missing hour late or skip it; run the repeated hour once), and the scheduler stores the next run as a UTC instant. That policy is a cost you can name.

(c) A probe of the core invariant. Keep the design and show why it holds: "Both claims target the same job row. The second waits, re-checks the lease, updates nothing and gets no fence."

Move 5: Re-derive out loud when the facts change

Pressure: 0:30. "Actually, each team may have at most 200 runs going at once."

The weak responses are panic ("Oh… then we'd need a different database?") and pretending nothing changed. The strong one has a shape: say what depends on the new fact, re-derive the numbers, update the board.

"That changes the claim, not the schedule. What depends on it: claiming (it now needs a team slot too), finishing and lease expiry (both must give the slot back). What doesn't: the scheduler, the queues and the history, as long as each team's limit covers its load, which is the next thing to check.

The limit is a count: one row per team with running and a maximum. A claim takes a slot with a conditional increment, adding 1 only while running is below the maximum, in the same transaction as the job's lease. Both rows live on the one primary, which is what makes that a single transaction. Runs that don't get a slot stay queued for their team instead of retrying in a tight loop.

Numbers: say the largest team owns 10% of runs, 200,000 a day. Over the whole day that's about 420 runs going at once on average (200,000 ÷ 86,400 × 180 s), so a limit of 200 is below its average load: 200 slots of 3-minute runs start about 67 a minute, 4,000 an hour, while 8,333 an hour arrive. Its backlog grows by about 4,300 an hour and never drains. Starting its top-of-hour runs within 60 seconds would take about 3,600 slots (3,333 due plus about 250 already running). Which do you want for that team: a higher limit, or rejecting its extra runs?"

That number comes from Little's law, covered in Key Concepts & Terminology: the average number of runs in progress equals their start rate times how long each lasts, L = λW. Here λ is 200,000 a day (about 2.3 a second) and W is 180 seconds, so L is about 420. Checked the other way, a limit fixes L at 200, so the most the team can start is λ = L ÷ W, about 67 a minute. Always check a limit against the whole day's load first, then against the burst.

Is one counter row a hot spot? Each update keeps the row locked until its transaction commits. With the synchronous-standby round trip, say 2 ms per transaction, one row tops out around 500 updates a second. With the limit binding, this row sees about one update per start and one per finish, roughly 2 a second; even with no limit it would see at most that team's 56 claims a second at the top of the hour, plus its finishes. One row is fine. Near the ceiling, you'd split the count across a few rows.

Then change the board: a new line in the numbers column ("team slots: 1 row per team; largest team needs about 420 on average, about 3,600 at hh:00") and the slot step on the run path. Here is the mechanism:

import sqlite3

db = sqlite3.connect(":memory:", isolation_level=None)
db.execute("""CREATE TABLE teams (team TEXT PRIMARY KEY,
              running INTEGER NOT NULL, max_running INTEGER NOT NULL)""")
db.execute("INSERT INTO teams VALUES ('payments', 198, 200)")

TAKE_SLOT = """UPDATE teams SET running = running + 1
               WHERE team = :team AND running < max_running"""
GIVE_BACK = "UPDATE teams SET running = running - 1 WHERE team = :team AND running > 0"

def take_slot(team):
    return db.execute(TAKE_SLOT, {"team": team}).rowcount == 1

print(take_slot("payments"), take_slot("payments"), take_slot("payments"))   # True True False
print(db.execute("SELECT running FROM teams").fetchone()[0])                  # 200: never above
db.execute(GIVE_BACK, {"team": "payments"})                                   # a run finished
print(take_slot("payments"))                                                  # True

Why it works: the check and the increment are one statement on one row, so there's no moment where two workers both see "199 running". In PostgreSQL the second writer waits, then re-checks running < max_running against the committed value, exactly as with the lease.

⚠️ The version that looks equivalent isn't: SELECT running, check it in application code, then an UPDATE that adds 1 with no condition. Two workers can both read 199, both pass the check and both add, and the team runs 201. The condition has to be in the write.

When you were the one who was wrong

Sometimes you catch your own mistake. Own it in one sentence and fix it: "Correction: I said one row could take thousands of updates a second, but the lock is held through the commit round trip, so it's hundreds. We need about 2, so the decision stands." No long apology, and no quiet redrawing in the hope that nobody noticed. Interviewers notice.

Your turn: 0:37, in the wrap-up. "What changes at 10×? Twenty million runs a day, still 40% at the top of the hour." What changes, what doesn't, and what are the new numbers?

Check your answer

The top of the hour now brings 333,333 due runs, 5,556 claims and as many trigger records a second, then up to 11,111 renewals and about 1,850 finishes a second: about 20,000 writes a second at the worst moment, if every run gets to start on time. That's past what I'd put on one primary, so the jobs get partitioned across several. The team limit decides the key: a claim and its team slot must stay in one transaction, so partition by team, not by job.

The price is uneven shards, and the limits decide how uneven. A team with 10% of runs would bring about 10% of that load to its shard, about 2,000 writes a second, if its limit let all its runs start on time; held to 200 slots, it would bring about a dozen a second and fall ever further behind. The limits have to scale with the load: that team now needs about 4,200 runs going at once on average and about 36,000 to start its top-of-hour runs on time. So every team's limit gets re-derived from its own daily load, and the shard plan from the limits. The queues carry 5,556 messages a second, which is still modest for a queue, and the claim logic doesn't change at all.

Move 6: Score it like an interviewer

Pressure: the mock ends. "That went OK, I think." A feeling can't tell you what to fix.

Score five rows from 1 to 4. This is a practice rubric, not any company's hiring scorecard (companies use their own); it covers what this course has asked you to do.

Row 1 2 3 4
Scope and numbers No questions, no numbers, or more than 10 minutes of questions Numbers that never change a decision Key requirements and two or three deciding numbers by 0:10 …and names the number that dominates the design
Working design No traceable path by the end A traceable path, but late (after about 0:25) or missing a core requirement One write and one read traced end to end by 0:20, with API and data model …and a board kept current as the design changes
Depth No deep dive on the riskiest part, or false claims A generic deep dive ("add a cache") The riskiest part with a correct mechanism and its failure modes, finished by 0:35 …and the wrap-up says what changes at 10×
Trade-offs Choices without reasons Reasons, but no costs Four parts for the major decisions …with costs and switch points tied to numbers
Pushback and change Caves or digs in: changes or keeps the design without reasons Keeps or changes with a vague reason, or checks only some of the pushback Checks the board; keeps or changes with reasons …and re-derives numbers and redraws quickly

Every score needs evidence: a timestamp and what was said or drawn. "Trade-offs: 3" is an opinion. "Trade-offs: 3, four parts for the lease and for replication, but at 0:12 the queues appeared with no reason" is something you can fix.

Scoring the two runs you've seen:

Row The opening story's candidate The traced run
Scope and numbers 1: from 0:00 to 0:13, twenty-one questions and no numbers 4: at 0:07, "the average is a trap; about 2,000 writes a second at the worst moment"
Working design 1: no traceable path from "job due" to "run started" at 0:40 4: path complete at 0:19 with API and tables, redrawn as the deep dive and the team slots arrived
Depth 1: no deep dive on the riskiest part; on the burst, only "we scale horizontally" at 0:27 4: the conditional claim, fencing and failover by 0:35; 10× in the wrap-up
Trade-offs 1: one reason in 40 minutes 3: the queues appeared at 0:12 with no reason; it came only when the interviewer asked at 0:24
Pushback and change 1: erased Cassandra at 0:34 with no reason 4: kept Postgres with numbers, added fencing, re-derived the limit at 0:30

Even a strong run has a lowest row. That's the point: the next mock targets it.

The 15-minute debrief

  1. Timeline. Write down the clock time you hit each checkpoint (0:10, 0:20, 0:35) next to the plan.
  2. Decisions. List every decision you said aloud, and mark which ones got all four parts.
  3. Interjections. For each one: what it was (hint, probe, new fact, failure, time), what you did, what you'd do now.
  4. Score the five rows, each with evidence.
  5. One micro-goal for the next mock: the lowest row, written as a behavior you can check. "Every box I draw gets a reason in the same breath", not "better trade-offs". One, not five.

Keep a log so you can see the trend:

Mock Prompt Path complete at Lowest row Micro-goal for next time
1 Job scheduler 0:26 Working design (2) Path traced end to end by 0:20; stop Scope at 0:10 to get there
2 File sync 0:19 Trade-offs (2) Every box gets a reason in the same breath
3 Nearby restaurants 0:18 Pushback and change (2) Check the board before answering any "why not"

🧠 A mock without a debrief mostly rehearses the habits you already have. The micro-goal is what makes mock 3 different from mock 2.

Your turn: debrief notes from a mock on a chat app. Requirements done at 0:09; path complete at 0:21; a deep dive on message ordering; wrap-up at 0:35. Seven decisions, six of them in four parts. At 0:24 the interviewer asked "why not deliver through a queue per user?"; you said "sure, that works too" and redrew the design around it. At 0:33, "what if a device is offline for a week?" got "we'd handle that" and a move to the next topic. Which row is lowest, and what's the micro-goal?

Check your answer

Pushback and change, at 1: at 0:24 you switched designs without comparing anything (caving), and at 0:33 you waved a failure question away. The trade-offs are fine, and a path at 0:21 is one minute late, close enough. A checkable micro-goal: "For every 'why not' or 'what if', restate it, check it against the board, and say keep or change with one reason before moving on."

Final round: three prompt cards

Real prompts don't say which move they need. Run each card as a 45-minute mock on the course's clock. Short on time? Run only Scope and Sketch and stop at the 0:20 checkpoint. If a friend holds the card, they reveal the hidden facts; if you're alone, open them only at the moment you'd have asked. Then open "What a strong run decides" and score yourself with the rubric.

Card 1: a folder on every device

"Design a file sync service: a folder that stays the same on all of a user's devices."

Hidden facts
  • 20 million daily active users with 2.5 devices each. An active user changes 10 files a day; a changed file averages 2 MB, and files can be up to 20 GB.
  • A change should reach the user's other online devices within 10 seconds.
  • Edits made on two devices must never silently overwrite each other.
  • Out of scope: sharing with other people, search, previews.
  • Planned pushback: "Why not store the files in the database?" and "Two devices edit the same file while offline. Who wins?"
What a strong run decides
  • Bytes, not metadata, dominate. 200 million changes a day is about 2,300 metadata commits a second (6,900 at a 3× peak), but 200 million × 2 MB is 400 TB a day uploaded, about 4.6 GB/s or 37 Gbit/s on average. File contents go in chunks (say 4 MB, named by their hash so an identical chunk is stored once) straight to object storage through pre-signed URLs, never through the API servers. The database holds small rows: each file's version and ordered chunk list. That's also the answer to "why not the database?".
  • Upload, then commit. A device uploads its chunks first, then commits the new chunk list; the server accepts the commit only if every listed chunk exists, so no other device is sent a file with missing pieces.
  • Conflicts are a conditional write. A commit names the version it started from, and the metadata update applies only if the file is still at that version. If another device committed first, the loser is saved as a "conflicted copy" beside the file. Cost: users merge those by hand; nothing is silently lost.
  • Notifications are a hint, not the truth. Online devices (up to 50 million) keep a long-poll or WebSocket connection; a change sends "your folder changed", and the device fetches the metadata since its last cursor. If notifications fail, devices fall back to polling every few minutes: freshness degrades, correctness doesn't. Degrading like this is safe, because the metadata database stays the source of truth; nothing here guards anything, so it isn't fail open or fail closed.

Card 2: restaurants near me

"Design the 'restaurants near me' search for a food app."

Hidden facts
  • 2 million restaurants worldwide; a few thousand change each day (openings, closures, hours). Changes may take a few minutes to show.
  • 50 million daily users, 5 searches each per day; the peak is 3× the average.
  • Results within 5 km, nearest first, filterable by "open now" and cuisine; p99 under 200 ms.
  • Planned pushback: "Why not just PostGIS?" and "What about a downtown with 5,000 restaurants within 1 km?"
What a strong run decides
  • Numbers: 250 million searches a day is about 2,900 a second, 8,700 at peak. The data is small: 2 million restaurants at about 1 KB each is 2 GB, which fits in memory on every replica. This is a read-scaling problem, not a storage problem.
  • A spatial index: geohash cells, a quadtree, or PostgreSQL with PostGIS and a GiST index. "PostGIS is a fine answer at 2 GB" is a strong reply to the pushback, as long as you add how reads scale: replicas behind the search API.
  • Cell edges: a restaurant 50 m away can sit in the neighboring cell, so search the cell and its neighbors, then filter by true distance.
  • Dense areas: cap the results per query and page through them, or use adaptive cells (a quadtree splits crowded cells).
  • Freshness: "open now" is computed at query time from opening hours; closures arrive as updates and show within minutes, which the requirements allow.

Card 3: the last ten seconds of an auction

"Design the bidding service for an online auction site."

Hidden facts
  • 1 million live auctions; most bids arrive in an auction's final minute.
  • A hot auction gets 100 bids a second in its last 10 seconds, with 20,000 people watching the price.
  • A bid is accepted only if it beats the current price by the minimum increment. The auction ends at a fixed time by the server's clock.
  • Watchers must see a new price within 1 second.
  • Planned pushback: "Two equal bids arrive a millisecond apart. Who wins?" and "Isn't the auction row a hot row?"
What a strong run decides
  • The invariant is that accepted bids strictly increase, and it lives in one conditional write per auction: accept only if the bid is at least the price plus the increment and the database's clock says the auction is still open.
  • Equal bids: the first to commit wins; the second fails the condition, because it no longer beats the price by the increment. The order is the database's commit order, not the bidders' clocks.
  • The hot row: 100 bids a second against a ceiling of a few hundred on one row (Move 5's arithmetic) fits. Near the ceiling, the next step is a single owner per auction that applies bids in memory, in order, and persists each result, which brings its own failover cost.
  • Price updates: "within 1 second" rules out lazy polling, because a page that polls once a second, reading a copy cached for up to a second, can show a price almost 2 seconds old. Push each new price over server-sent events or WebSockets; 20,000 open connections for a hot auction is a modest connection tier. Compare the scheduler, where a second of delay inside a 60-second window made polling the database the simple choice.

Cheat sheet: pressure → move → cost

When… Reach for What it costs
It's 0:20 and there's no full path Compress, don't skip: assume aloud, draw the simplest path Some requirements are assumed, not confirmed
You name a component Decision, because, cost, switch if About 20 seconds per decision
Boxes pile up Request paths with labeled arrows; numbers on the left Slower drawing
"Why not X?" Restate, check the board, decide, name the cost Sometimes admitting that X is better
A hint Take it now, and say so Your current thread stays unfinished
A fact changes Say what depends on it; re-derive; redraw A minute of rework, out loud
A dependency dies Fail open or closed, chosen from the requirement Lost protection, or a paused service
The mock ends Five rows with timestamps; one micro-goal Fifteen minutes of debrief

Before moving on, run one timed mock this week: a card above, or a design from this course you haven't yet practiced aloud, such as the URL shortener, the rate limiter, the news feed, video streaming or chat. Record it. Then, from a blank board, explain aloud as you would in the wrap-up at 0:35: the three decisions that mattered, the requirement each one served, what each one cost, and what would make you change it. If that takes you two minutes and no notes, you're ready for the real room.

For the rest of final prep (consistent hashing, consensus, sagas and a checklist), see Advanced Topics & Final Prep.