fivenines
2/38

Theory Tutorial

Non-functional requirements: nines math and the scale envelope

Time
5m
Level
not specified
Artifacts
theory + practice
Progress0%
Lesson 0.3

Non-functional requirements: nines math and the scale envelope

From 0.2: the exchange core (presence → matching → trips) is the latency- and availability-critical loop; settlement and analytics are not. Keep that asymmetry in mind; availability arithmetic now gives the distinction a budget.

A first pass might give every service the target availability and assume the end-to-end flow inherits the same number. That works only while the failure case is absent. An availability claim must be computed over the complete user flow, including serial dependencies, correlated failures, recovery time, and the actual scale envelope.

What a nine actually is

Availability = successful minutes ÷ total minutes, over a window (we use 30 days). Each extra nine cuts the allowed failure by 10×:

AvailabilityDowntime / yearDowntime / 30 daysEngineering posture
99.9% Part 18 h 46 m43 m 50 sHumans can react. Redundant hardware, manual failover, on-call fixes things before the budget burns.
99.99% Part 252 m 36 s4 m 23 sHumans are too slow for detection + failover. Recovery must be automated; failures must be contained.
99.999% Part 35 m 15 s26 sEven automated failover is too slow if it's reactive. The system must be statically stable: pre-provisioned, active-active, degrading instead of failing.

The multiplication trap

Availabilities of serial dependencies multiply. A request that must touch five services, each individually 99.9%, is a 99.5% request — over 3.6 hours of failure a month. This single fact drives most of Parts 2 and 3: shorten critical paths, make dependencies asynchronous or optional, and add redundancy (parallel paths multiply unavailabilities instead: two independent 99.9% paths ≈ 99.9999%).

flowchart TB
  subgraph S ["Serial chain — availabilities multiply"]
    A1["Gateway 99.9%"] --> A2["Matching 99.9%"] --> A3["Location index 99.9%"] --> A4["Trip store 99.9%"] --> A5["Pricing 99.9%"]
  end
  S --> SR["End-to-end: 0.999^5 ≈ 99.5% — about 3.6 h down per month"]
  subgraph P ["Redundant pair — unavailabilities multiply"]
    B1["Path A 99.9%"]
    B2["Path B 99.9%"]
  end
  P --> PR["Either path suffices: 1 − 0.001² ≈ 99.9999%"]
Why deep synchronous call chains are the enemy of availability, and why redundancy is the cure — if the paths fail independently.

The scale envelope

All capacity discussions in this course use these numbers. Learn them once (this is CLT pre-training — isolating the numeric facts now so later lessons don't have to teach them mid-argument):

DimensionFigureDerived load
Riders50 M monthly active~5 M daily active
Drivers3 M registered1 M online at global peak
Location pings1 every 4 s per online driver250 K writes/s peak; ~2 TB/day raw
Trips20 M/day~1,000 requests/s peak (dispatch)
Fare quotes~5 quotes per request~5,000/s peak
Payments20 M captures/day~700/s peak; every cent double-entry ledgered
Cities~700, across ~6 geographic regionsnatural sharding and cell boundary

Availability targets are meaningless without a definition of "down." We define per-flow SLIs in 2.1 — e.g. "dispatch is up if ≥ 99% of ride requests get a driver offer within 10 s." A green server dashboard while riders stare at spinners is still down.

Next step

See what actually stuck.

Take the practice scenarios now.