fivenines
2/38
Lesson 0.3

Non-functional requirements: nines math and the scale envelope

🎯 Objectives
  • Translate 99.9 / 99.99 / 99.999% into allowed downtime and into engineering posture.
  • Compute why chained dependencies destroy availability, and memorize the scale envelope used in every later lesson.
🔗 Connect

From 0.2: the exchange core (presence → matching → trips) is the latency- and availability-critical loop; settlement and analytics are not. Keep that asymmetry in mind — this lesson gives it numbers.

What a nine actually is

Availability = successful minutes ÷ total minutes, over a window (we use 30 days). Each extra nine cuts the allowed failure by 10×:

AvailabilityDowntime / yearDowntime / 30 daysEngineering posture
99.9% Part 18 h 46 m43 m 50 sHumans can react. Redundant hardware, manual failover, on-call fixes things before the budget burns.
99.99% Part 252 m 36 s4 m 23 sHumans are too slow for detection + failover. Recovery must be automated; failures must be contained.
99.999% Part 35 m 15 s26 sEven automated failover is too slow if it's reactive. The system must be statically stable: pre-provisioned, active-active, degrading instead of failing.

The multiplication trap

Availabilities of serial dependencies multiply. A request that must touch five services, each individually 99.9%, is a 99.5% request — over 3.6 hours of failure a month. This single fact drives most of Parts 2 and 3: shorten critical paths, make dependencies asynchronous or optional, and add redundancy (parallel paths multiply unavailabilities instead: two independent 99.9% paths ≈ 99.9999%).

flowchart TB
  subgraph S ["Serial chain — availabilities multiply"]
    A1["Gateway 99.9%"] --> A2["Matching 99.9%"] --> A3["Location index 99.9%"] --> A4["Trip store 99.9%"] --> A5["Pricing 99.9%"]
  end
  S --> SR["End-to-end: 0.999^5 ≈ 99.5% — about 3.6 h down per month"]
  subgraph P ["Redundant pair — unavailabilities multiply"]
    B1["Path A 99.9%"]
    B2["Path B 99.9%"]
  end
  P --> PR["Either path suffices: 1 − 0.001² ≈ 99.9999%"]
Why deep synchronous call chains are the enemy of availability, and why redundancy is the cure — if the paths fail independently.

The scale envelope

All capacity discussions in this course use these numbers. Learn them once (this is CLT pre-training — isolating the numeric facts now so later lessons don't have to teach them mid-argument):

DimensionFigureDerived load
Riders50 M monthly active~5 M daily active
Drivers3 M registered1 M online at global peak
Location pings1 every 4 s per online driver250 K writes/s peak; ~2 TB/day raw
Trips20 M/day~1,000 requests/s peak (dispatch)
Fare quotes~5 quotes per request~5,000/s peak
Payments20 M captures/day~700/s peak; every cent double-entry ledgered
Cities~700, across ~6 geographic regionsnatural sharding and cell boundary
⚠️ Pitfall

Availability targets are meaningless without a definition of "down." We define per-flow SLIs in 2.1 — e.g. "dispatch is up if ≥ 99% of ride requests get a driver offer within 10 s." A green server dashboard while riders stare at spinners is still down.

Next step

See what actually stuck.

Take the practice scenarios now.