How to Pass a System Design Interview in 45 Minutes
Quick summary
- Lead with the problem. Confirm the core use cases and constraints before you draw boxes.
- Use estimates to make choices. Round numbers are enough when they show what might become a bottleneck.
- Start with a small design. Trace one complete request before adding caches, queues, replicas, or shards.
- Deep-dive where the requirements hurt. Explain the trade-off, the failure mode, and why your choice fits.
- Finish with operations. Reliability needs measurable goals, bounded failure handling, and signals an operator can act on.
A system design interview can begin with three words: “Design a ride-hailing app.” You have an empty board, a short clock, and a prompt broad enough to swallow the whole product.
The easy mistake is to start drawing infrastructure. The stronger move is to take control of the conversation. Turn the vague prompt into a small problem, make your assumptions visible, and build only as much architecture as those assumptions justify.
This five-stage system design interview framework gives you a way to do that in 45 minutes. Treat the times as guardrails, not a script. The interviewer may steer you toward a different area, and you should follow the most important signal.
The 45-minute system design interview at a glance
| Time | Phase | What you need by the end |
|---|---|---|
| 0 to 5 minutes | Scope and requirements | A short list of core use cases, constraints, and non-goals |
| 5 to 10 minutes | Estimates and contracts | Order-of-magnitude traffic, key APIs, and data access patterns |
| 10 to 20 minutes | High-level architecture | A simple diagram and a narrated end-to-end request |
| 20 to 35 minutes | Deep dives and bottlenecks | A reasoned answer to the hardest constraint or likely bottleneck |
| 35 to 45 minutes | Resilience and observability | Failure behavior, measurable service goals, and a concise summary |
The source 45-minute blueprint uses this same five-part rhythm. The goal here is to turn that rhythm into a practical conversation, while keeping the design grounded in the requirements you uncover.
Phase 1: Scope the system design interview problem (0 to 5 minutes)
Suppose the interviewer says, “Design Uber.” Do not start with a database schema. First ask what part of the product matters in this interview.
For a basic ride-hailing flow, you might agree on three core features:
- A rider requests a ride.
- A driver accepts or rejects a ride.
- The rider can see the driver’s location while the trip is active.
Then clarify constraints that could change the design. Do bookings need a durable confirmation? How fresh must driver locations be? What region and peak traffic should we assume? Is this a global service, or are we focusing on one city?

Be precise about reliability language. Ask for a user-facing latency or availability goal if one matters. A service-level objective, or SLO, is an internal target for a measured service outcome. A service-level agreement, or SLA, is a commitment made to customers, often with contractual consequences. They are related, but they are not interchangeable. Google’s SRE guidance explains how objectives should reflect what users need and the cost of achieving more reliability.
Write down the answer and name a non-goal. “I’ll focus on requesting, matching, and live trip location. I’ll leave payments and ratings out unless we have time.” That single sentence protects the rest of the interview from scope creep.
Rule: Do not design for a scale the prompt has not established.
Phase 2: Estimates and API contracts in a system design interview (5 to 10 minutes)
Your arithmetic does not need false precision. It needs to reveal which resources could matter.

For an illustrative example, assume 10 million daily users each generate 10 relevant events. That is 100 million events per day, or about 1,200 events per second on average. If we assume a three-times peak, we should reason about roughly 3,600 events per second at peak. Those are assumptions, not universal traffic facts. Ask the interviewer whether the scale is in the right range.
If each event averages 10 KB, peak ingress would be about 36 MB per second before protocol overhead, replication, or any extra copies. That number does not prove we need a particular server or database. It tells us bandwidth and write throughput deserve attention, and gives us a load to benchmark against candidate components.
Next, make the system’s contracts concrete. A ride request could use a synchronous HTTPS API that returns a ride ID and current status. For location updates, WebSockets fit when both sides need to send messages over a live connection; server-sent events fit one-way updates from server to client; periodic polling may be adequate when freshness requirements are looser.
Match the data store to its access pattern. A relational store is a natural candidate for durable ride and payment records with transactional constraints. A geospatial index or in-memory store may suit nearby-driver lookups and rapidly changing coordinates. These are starting hypotheses, not automatic technology choices. State what you need the store to do before naming a product.
Rule: Every estimate should support a decision. If the number changes nothing, move on.
Phase 3: Draw the smallest useful architecture (10 to 20 minutes)
Start with a few components: client, load balancer, API gateway, ride service, and storage. Add a location service only if it clarifies the core flow. Keep the picture small enough that you can explain every box.

Then trace one request from beginning to end: “The rider sends a request. The gateway authenticates it and applies rate limits. The ride service creates a pending ride and returns its ID. A matching process finds candidate drivers, and the chosen driver accepts. The service updates the ride state and notifies both clients.”
That narration exposes missing decisions. Is matching part of the synchronous response or an asynchronous step? Can a driver accept the same ride twice? What does the rider see if matching takes too long? You can now answer those questions in context rather than adding disconnected components.
Give every box a job. A load balancer might spread requests across stateless service instances and remove unhealthy instances from rotation. A queue might separate matching work from the API response. A cache might hold recently updated driver locations. If you cannot explain the job, remove the box.
For another example of how an Uber-related design changes with its goal, see this White Way Web article on demand prediction and driver repositioning. A system that predicts demand has different data flows and trade-offs from one that processes an individual ride request.
Rule: Prove the happy path works before you scale it.
Phase 4: Bottlenecks in a system design interview (20 to 35 minutes)
Now invite pressure. Ask what might fail first at the agreed scale, or let the interviewer choose a constraint. Pick one or two areas and work through them. Trying to solve every distributed-systems problem at once is a reliable way to finish none of them.
When reads dominate
If clients read driver locations far more often than drivers update them, a short-lived cache may reduce pressure on the location store. Explain the freshness trade-off: cached coordinates can be slightly stale. Set an expiry based on the user experience you agreed on, and describe what happens on a cache miss. A cache helps only if its update and expiry behavior fits the data.
When writes spike
A queue can absorb bursts and let workers process updates at a controlled rate. It does not make excess work disappear. If events arrive faster than workers can handle them for long enough, the backlog grows. Set limits, monitor queue age, and decide whether to slow producers, reject low-priority work, or shed load.
Partitioning and hot keys

Partitioning can increase write capacity, but the partition key matters. A single city key could concentrate a busy city on one partition. Hashing every driver may spread writes, but can make nearby-driver queries harder. Explain which operation you are optimizing and how the design handles the query that the key makes less convenient.
When two actions race
Ride acceptance is a concurrency problem: two drivers might attempt to claim one ride at nearly the same time. Make the state transition conditional so only one acceptance succeeds. For retried requests, use an idempotency key or equivalent deduplication mechanism so a retry does not create a second ride.
When regions or dependencies fail
Do not label an entire design “AP” or “CP” and move on. During a network partition, consistency and availability can conflict, but the useful interview question is narrower: for this operation, what must still work, and what stale or rejected result is acceptable? A location marker may tolerate brief staleness. A ride assignment may need one authoritative winner.
Rule: Name the requirement, compare the options, and own the downside of your choice.
Phase 5: Make the design operable (35 to 45 minutes)

A design is unfinished if nobody can tell when it is failing. Pick a few signals tied to the user journey: ride-request success rate, matching latency, location freshness, and the age of work waiting in a queue. If you mention latency, say whether you care about the typical request or the slow tail, such as the 95th or 99th percentile.
Then explain how components behave under stress. Set timeouts so requests do not wait forever. Retry only operations that are safe to repeat, cap the attempts, and use exponential backoff with jitter to avoid synchronized retry bursts. AWS’s Builders’ Library describes why retries can amplify an outage and how idempotent APIs make retry behavior safer. A circuit breaker can stop repeated calls to a failing dependency, but it needs a recovery path and a clear fallback.
Finally, state what failover means. A standby region is not magic. Discuss how much recent data could be lost, how long recovery may take, and how you prevent two regions from both accepting conflicting writes. Google’s SRE material frames reliability around explicit objectives and error budgets rather than a promise of perfect uptime.
Close with a short synthesis: “We kept ride creation transactional, separated high-volume location updates, and used a queue for matching bursts. The main trade-off is location freshness versus read load. I would next validate the peak write estimate and test matching behavior during a worker outage.”
Rule: End with the risk you would investigate next, not another technology name.
Three habits that weaken a system design interview answer
- Overbuilding before the requirements are clear. Multi-region sharding is not a default opening move.
- Defending a design instead of examining it. Treat interviewer pushback as a chance to test assumptions and compare options.
- Spending the whole clock on one phase. A perfect estimate cannot make up for a design you never finish explaining.
What a strong system design interview answer demonstrates
A strong answer is not the one with the most boxes. It is the one where each choice follows from a stated need, each important flow can be traced, and the candidate can explain what happens when an assumption fails.
Use the five phases to keep your footing: scope the problem, estimate what matters, draw a simple baseline, deepen the hard parts, and finish with failure behavior and observability. The structure gives you room to show judgment. The judgment is still yours.