Non-functional requirements: nines math and the scale envelope
- Translate 99.9 / 99.99 / 99.999% into allowed downtime and into engineering posture.
- Compute why chained dependencies destroy availability, and memorize the scale envelope used in every later lesson.
From 0.2: the exchange core (presence → matching → trips) is the latency- and availability-critical loop; settlement and analytics are not. Keep that asymmetry in mind — this lesson gives it numbers.
What a nine actually is
Availability = successful minutes ÷ total minutes, over a window (we use 30 days). Each extra nine cuts the allowed failure by 10×:
| Availability | Downtime / year | Downtime / 30 days | Engineering posture |
|---|---|---|---|
| 99.9% Part 1 | 8 h 46 m | 43 m 50 s | Humans can react. Redundant hardware, manual failover, on-call fixes things before the budget burns. |
| 99.99% Part 2 | 52 m 36 s | 4 m 23 s | Humans are too slow for detection + failover. Recovery must be automated; failures must be contained. |
| 99.999% Part 3 | 5 m 15 s | 26 s | Even automated failover is too slow if it's reactive. The system must be statically stable: pre-provisioned, active-active, degrading instead of failing. |
The multiplication trap
Availabilities of serial dependencies multiply. A request that must touch five services, each individually 99.9%, is a 99.5% request — over 3.6 hours of failure a month. This single fact drives most of Parts 2 and 3: shorten critical paths, make dependencies asynchronous or optional, and add redundancy (parallel paths multiply unavailabilities instead: two independent 99.9% paths ≈ 99.9999%).
flowchart TB
subgraph S ["Serial chain — availabilities multiply"]
A1["Gateway 99.9%"] --> A2["Matching 99.9%"] --> A3["Location index 99.9%"] --> A4["Trip store 99.9%"] --> A5["Pricing 99.9%"]
end
S --> SR["End-to-end: 0.999^5 ≈ 99.5% — about 3.6 h down per month"]
subgraph P ["Redundant pair — unavailabilities multiply"]
B1["Path A 99.9%"]
B2["Path B 99.9%"]
end
P --> PR["Either path suffices: 1 − 0.001² ≈ 99.9999%"]
The scale envelope
All capacity discussions in this course use these numbers. Learn them once (this is CLT pre-training — isolating the numeric facts now so later lessons don't have to teach them mid-argument):
| Dimension | Figure | Derived load |
|---|---|---|
| Riders | 50 M monthly active | ~5 M daily active |
| Drivers | 3 M registered | 1 M online at global peak |
| Location pings | 1 every 4 s per online driver | 250 K writes/s peak; ~2 TB/day raw |
| Trips | 20 M/day | ~1,000 requests/s peak (dispatch) |
| Fare quotes | ~5 quotes per request | ~5,000/s peak |
| Payments | 20 M captures/day | ~700/s peak; every cent double-entry ledgered |
| Cities | ~700, across ~6 geographic regions | natural sharding and cell boundary |
Availability targets are meaningless without a definition of "down." We define per-flow SLIs in 2.1 — e.g. "dispatch is up if ≥ 99% of ride requests get a driver offer within 10 s." A green server dashboard while riders stare at spinners is still down.