System Design: Estimation and Capacity Planning

Master back-of-the-envelope estimation, Jeff Dean's latency scale, QPS sizing, 5-year storage projections, ingress/egress bandwidth formulas, Pareto 80/20 cache sizing, and traffic spike headroom planning.

1. Back-of-the-Envelope Estimation & Mental Math Shortcuts

Back-of-the-envelope calculations determine whether a system fits on a single machine or requires a distributed cluster. Using rounded constants (86,400s ≈ 100,000s) allows rapid mental estimation during system design interviews and capacity planning.

Estimation Mental Math Cheat Sheet (QPS Conversions & Data Scale Powers)

System Design

Converting daily request volumes to QPS and navigating Byte, KB, MB, GB, TB, and PB powers of 10

100%
Loading system design canvas…

QPS Mental Math Shortcuts

Rapid conversion from daily requests to QPS.

  • 1 Day = 86,400 seconds (Round to 100,000 for quick interview mental math)
  • 1 Million requests/day = ~12 QPS
  • 100 Million requests/day = ~1,200 QPS
  • 1 Billion requests/day = ~12,000 QPS

Data Size Powers of 10

Standard storage unit scales.

  • 1 KB = 10^3 Bytes (Tiny payload, small JSON)
  • 1 MB = 10^6 Bytes (Image, compressed media)
  • 1 GB = 10^9 Bytes (RAM capacity, table partition)
  • 1 TB = 10^12 Bytes (Database volume)
  • 1 PB = 10^15 Bytes (Multi-year Big Data lake)

2. Latency Numbers Every Programmer Should Know (Jeff Dean Scale)

Understanding the physical speed differences between CPU caches, RAM, SSDs, HDDs, and network hops explains why caching and in-memory architectures are essential for low-latency systems.

Latency Scale Hierarchy (CPU Cache to Cross-Continental Network)

System Design

Visualizing physical hardware access times from 0.5 nanoseconds to 150 milliseconds

100%
Loading system design canvas…

Hardware Latency Insights in System Design

Architectural decisions dictated by hardware physics.

  • RAM vs SSD: Reading from RAM (100ns) is ~1,000x faster than reading from an NVMe SSD (100μs)
  • SSD vs HDD: NVMe SSD random reads (100μs) are ~100x faster than spinning mechanical HDD seeks (10ms)
  • Network vs Local: A cross-continental network roundtrip (150ms) is 1.5 million times slower than a local RAM read

3. Estimating QPS (Queries Per Second) & Peak Load

Estimating traffic throughput begins by identifying Daily Active Users (DAU) and breaking down user behavior into read and write frequencies.

End-to-End QPS & Throughput Sizing Pipeline Flowchart

System Design

Calculating Average QPS, Peak QPS (2.5x), and determining the Read-to-Write architecture ratio

100%
Rendering diagram…

Step-by-Step QPS Calculation Example (Twitter / Social Scale)

Assumptions: 300 Million DAU, 10 reads/day, 1 write/day.

  • Total Daily Reads: 300M × 10 = 3 Billion reads/day
  • Average Read QPS: 3B / 86,400s ≈ 35,000 QPS
  • Peak Read QPS (2.5x multiplier): ~87,500 QPS
  • Total Daily Writes: 300M × 1 = 300 Million writes/day
  • Average Write QPS: 300M / 86,400s ≈ 3,500 QPS (Peak: ~8,750 QPS)

Read vs Write Ratio Architectural Decisions

Read:Write Ratio = 35,000 : 3,500 = 10:1 (Read-Heavy).

  • Read-Heavy (>10:1): Heavy caching (Redis/Memcached), Read Replicas, CDN Edge distribution
  • Write-Heavy (<5:1): LSM-Tree stores (Cassandra), Kafka write buffering, message batching

4. Estimating Storage & Bandwidth Requirements (5-Year Projections)

Storage sizing must distinguish between structured relational/document metadata (stored on fast SSDs with 3x replication and index overhead) and raw media blobs (images/videos stored on cheap object stores like AWS S3).

5-Year Storage Projections & Ingress/Egress Bandwidth Throughput

System Design

Factoring 3x replication, 20% database index overhead, and CDN egress offloading

100%
Rendering diagram…

5-Year Storage Sizing Formulas

Accounting for real-world production overheads.

  • Daily Metadata: 300M writes × 500 Bytes = 150 GB / day
  • 5-Year Raw Metadata: 150 GB × 365 × 5 = 274 TB
  • Total DB Disk (3x Replication + 20% Index): 274 TB × 3 × 1.2 ≈ 1 PB (SSD)
  • 5-Year Media Storage (10% has 200KB image): 6 TB/day × 365 × 5 ≈ 11 PB (S3)

Bandwidth Throughput Sizing

Network pipe requirements.

  • Ingress (Incoming Write): 6.15 TB / 86,400s = 71.2 MB/s (570 Mbps)
  • Egress (Outgoing Read): 712 MB/s (5.7 Gbps)
  • CDN Strategy: CloudFront/Cloudflare absorbs 95% of egress bandwidth at edge POPs

5. Estimating Cache Size (Pareto 80/20 Rule)

You do not need to cache all data in expensive RAM. According to the Pareto 80/20 Rule, 20% of items generate 80% of read traffic. Caching the daily active 20% working set delivers an optimal cache hit ratio.

Pareto 80/20 Cache Working Set Allocation & Redis Sizing Model

System Design

Calculating daily hot working set RAM requirements with safety overhead buffers

100%
Loading system design canvas…

Cache RAM Sizing Formula

Step-by-step memory allocation.

  • Daily Hot Working Set = 20% × Daily Metadata Volume = 0.20 × 150 GB = 30 GB
  • Safety Overhead Buffer (2x for key metadata, feed lists & Redis memory fragmentation) = 30 GB × 2 = 60 GB
  • Hardware Provisioning: 4x Redis cluster nodes with 32 GB RAM each = 128 GB Total RAM

Production System Design Capacity & Estimation Calculator in Node.js

A Node.js utility that estimates storage, bandwidth, and server capacity needs from traffic and data assumptions.

6. Peak Load, Flash Sales & Traffic Spikes Planning

Systems never experience uniform traffic. Viral events, breaking news, and flash sales trigger sudden 3x to 5x load spikes. Resilient systems provision adequate headroom and buffer write bursts.

Traffic Spike Headroom, Burst Buffers & Autoscaling Workflow

System Design

Managing viral traffic surges using Kafka queues, pre-warmed pod pools, and edge shedding

100%
Rendering diagram…

Traffic Spike Resilience Strategies

Surviving unexpected traffic surges.

  • Provisioning Headroom: Target steady-state CPU utilization at 50-60% to absorb instant spikes before autoscaling triggers
  • Message Queue Buffering: Buffer write spikes in Kafka/SQS to prevent database connection exhaustion
  • Graceful Degradation: Shed non-critical telemetry and recommendations during peak loads

Knowledge Check

1. What is the quick mental math shortcut for converting 100 Million daily requests to QPS?

2. According to Jeff Dean latency numbers, how much faster is RAM access than SSD read?

3. How does the 80/20 Pareto Principle apply to Redis cache sizing?

4. What is the difference between Ingress and Egress network bandwidth?

5. Why must 5-year database storage estimates include a 3x replication factor?