Core Concept

Back-of-the-Envelope Estimation

Interviewers don't expect exact numbers — they want quick math that ties user scale to QPS, storage, and bandwidth so your design choices rest on real constraints.


1. What It Is

Before you draw boxes on the whiteboard, spend two minutes on napkin math. Interviewers want to see that your cache, CDN, and shard decisions rest on real constraints — not vibes.

What:

Quick arithmetic you run on the whiteboard to approximate QPS, storage, memory, and bandwidth from a few assumptions (DAU, actions per user, payload size).

Primary purpose:

Ground your design choices — caches, CDNs, shards — in numbers that prove the architecture can handle the stated scale.

Usually used for:

Sizing cache clusters, planning database capacity, and estimating whether origin bandwidth can serve traffic without a CDN.

2. Core Mental Model

We keep three anchors in our back pocket — they cover most interview math without reaching for a calculator:

⏱️ The 86,400 Anchor

There are 86,400 seconds in a day. Divide daily volume by that to get average QPS — then multiply by 3–10x for peak load and state both numbers.

🗄️ Storage Durability (3x)

Raw data size is not your storage bill. Multiply by a 3x replication factor (or whatever RF your design uses) before comparing against disk limits.

📈 Headroom Safety Buffer

Add 30–50% headroom on top of peak estimates so a traffic spike does not immediately exhaust capacity.

In the room

Candidates often skip math entirely or spend ten minutes on precision. State your assumptions out loud, round aggressively, and move on once the order of magnitude is clear — the interviewer will redirect you if they want more detail.

3. Why It Matters in HLD

When the interviewer drops a scale number — 100M DAU, 1B messages per day — this is where we prove the design is not hand-waving. We frame the discussion around three lenses:

Needed When:

When the interviewer asks how many DB instances or how large the Redis cluster should be.

Avoids:

Designs that look plausible on a diagram but cannot survive real traffic, plus over-provisioned hardware and surprise scaling failures.

Optimizes For:

Decisions backed by numbers, balanced infrastructure spend, and clear justification when you add a cache, CDN, or shard.

4. Architecture & Data Flow

Walk the diagram left to right as interview steps you narrate out loud. Step 1 — Gather inputs: state DAU, actions per user per day, and average payload size. Step 2 — Average QPS: divide daily volume by 86,400 and round aggressively. Step 3 — Peak QPS: multiply by your peak factor (3–10x) and say both average and peak. Step 4 — Size the app tier: divide peak QPS by ~1,000 QPS per server. Step 5 — Storage and bandwidth: multiply peak QPS x bytes per request; convert Gbps to GB/s before comparing to disk or NIC limits. Step 6 — Add headroom: layer 30–50% buffer and call out the first component that breaks.

Loading...

In the room

If your storage estimate is 500 GB but you propose 50 Redis nodes, pause and sanity-check. Interviewers respect candidates who catch their own order-of-magnitude errors and revise out loud.

5. Key Characteristics

These anchor numbers are our whiteboard cheat sheet — we keep a few memorized so we never stall mid-interview:

  • The Core Numerical Anchors: Round numbers to keep whiteboard math fast:
MetricApproximation ValueSystem Design Use Case
Seconds in a Day86,400Direct QPS conversions
1 MB (Mega)1 Million BytesMedia asset payloads
1 GB (Giga)1 Billion BytesBulk database storage volumes
Standard Replication3x (Default)Storage overhead calculation
  • Modern Hardware Limits (Interview Cheat Sheet): Round numbers that anchor latency and capacity math on the whiteboard:
MetricApproximation ValueSystem Design Use Case
RAM Random Access~100 ns
  • In-memory cache lookups
  • L1/L2 cache hits are ~1–10 ns
SSD Sequential Read~500 MB/sDisk I/O budget for log replay, SSTable scans, cold storage reads
Same-Datacenter RTT~0.5–1 ms
  • Microservice hop budget
  • 10 chained RPCs ≈ 5–10 ms overhead
Cross-Region RTT~50–150 msJustifies async replication, regional read replicas, and CDN edge caching
Single App Server Throughput~1,000 QPS (typical)1M QPS ÷ 1K QPS/server ≈ 1,000 servers before caching and batching
Seconds in a Day86,400Convert daily volume to average QPS (always state peak multiplier separately)
1 Million Requests/Day~12 QPS avg
  • Small-scale sanity check: 1M / 86,400 ≈ 12
  • a single node handles this easily
1 Billion Requests/Day~11,574 QPS avg

6. Strategic Tradeoffs

Napkin math buys credibility, but it has a cost. We name the trade-off openly so the interviewer knows we are not cargo-culting formulas:

BenefitCost
Quantitative Grounding (proves your whiteboard architecture is feasible and sizing matches real physical limits)Estimation Delays (spending too long on simple arithmetic distracts from the core high-level design details)
Early Bottleneck Detection (QPS and bandwidth math surfaces CDN, cache, or shard needs before you over-build)False Precision (rounding errors compound — state assumptions explicitly and sanity-check against hardware anchors)
Interview Signal (numbers justify why you added Redis or rejected a single-server design)Scope Creep (deep-diving storage math on a 45-minute problem can crowd out API and consistency discussion)

7. Failure / Bottleneck Awareness

Interviewers love failure-mode probes on estimation. We volunteer these before they have to dig:

👾 Bit vs Byte Confusion

Problem: Network bandwidth is measured in bits per second (Gbps) while storage is in bytes (GB). Mixing the two gives an 8x error and under-provisions links.

Mitigation: Convert egress to bytes (divide by 8) before comparing against storage I/O or server throughput limits.

🛑 Designing for Average Load Only

Problem: Sizing for average QPS ignores peak hours — product launches, morning feed checks, or game releases can run 3–10x above average and take the system down.

Mitigation: Always compute peak QPS with a multiplier and size queues and autoscaling to absorb that burst.

🧩 Sharding Before Doing the Math

Problem: A large DAU number tempts candidates to propose dozens of shards immediately. A 10M-user app with 5 writes/user/day is only ~580 average write QPS — a single PostgreSQL primary with read replicas often handles that.

Mitigation: Compute QPS and storage first. Shard when a single node's write throughput, disk I/O, or RAM ceiling is provably exceeded — not because the user count sounds big.

8. Common HLD Usage

These are the whiteboard moments where a two-minute calculation changes the architecture — not just the slide deck:

Whiteboard ProblemEstimation Driving Decision
Twitter Timeline EgressProving that a 20 GB/s egress bandwidth cannot come from one server, forcing CDN + edge planning
YouTube Video UploadCalculating raw daily ingress bandwidth to justify a chunked, asynchronous transcoding queue
Distributed Cache SizingDetermining RAM size required to cache the top 20% hot-read keys (Pareto Principle)

9. Decision Signals

You will know it is time to pull out the envelope when the prompt sounds like one of these:

🎯 Think Estimation When:
  • The evaluator starts the interview with a scale metric (e.g., 'Design YouTube at 100M active daily users').
  • You must justify routing static media files away from primary databases to edge CDNs.
  • You need to decide whether a database table can fit completely inside a single Redis cache node's RAM.

11. Deep Dive (Optional)

Worked Example: Twitter Timeline Bandwidth (30s Speed Run)

Assumptions: 300M DAU, each user loads timeline 30x/day, returning 20 tweets per query. Average tweet JSON size = 2 KB.

Daily Reads = 300M * 30 = 9 Billion Reads/day
Average Read QPS = 9B / 86,400 = ~104,166 QPS
Peak QPS (5x multiplier) = ~520,833 QPS
Egress Bandwidth = 520,833 QPS * 20 tweets * 2 KB = 20.8 GB/s (or ~166 Gbps)

Design conclusion: ~20 GB/s egress cannot be served from database nodes directly. You need a geographically distributed CDN edge and precomputed Redis timelines to shield internal services.

💬Review

Help Us Improve

How helpful was this walkthrough?

Click a star to rate. We actively use this feedback to refine and update our system design content.

Placeholder
Optional but highly appreciated!

Discussion

Share your thoughts, ask questions, or help others.

Loading comments...