Lesson 1.14 99.9%
Recap: what 99.9% buys — and what breaks
🔗 Connect
This is a consolidation lesson — pure germane load. Nothing new: we replay Part 1 as one picture and one honest failure inventory, which becomes Part 2's to-do list.
You now have a complete ride exchange: onboarding through matching through money through reporting, running in one region across three zones, watched by symptom alerts and a human rotation. Its availability story: instance and zone failures self-heal; database failover takes minutes of human action; everything else is on-call plus runbooks inside a 43-minute monthly budget. That story is legitimate — and it has a ceiling. The honest failure inventory:
| Failure | At 99.9% today | Budget damage (30 d = 43 m 50 s) | Fixed in |
|---|---|---|---|
| Instance or zone loss | Automatic re-route, seconds | Negligible | — |
| Database primary dies | Human failover, 5–10 min | ~20% of budget per event | 2.3 |
| Bad deploy ships a bug | Rolling deploy spreads it everywhere; rollback after detection | 10–30 min — most of the budget | 2.8 |
| Downstream vendor (PSP, maps) browns out | Requests hang or fail; manual mitigation | Depends on their SLA — coupled fate | 2.6 |
| Traffic spike (concert, storm) | Autoscaling may lag; overload cascades | Unbounded | 2.6 |
| One service's failure cascades via sync calls | Nothing stops it | Unbounded | 2.4, 2.6 |
| Region failure | Total outage until region returns | Hours — years of budget | 2.7, 3.2 |
flowchart TB P1["Part 1 system — one region, humans in the loop"] --> W1["Self-healing: instances, zones"] P1 --> W2["Human-healing: DB failover, bad deploys, vendor issues"] P1 --> W3["No healing: region loss, cascades, overload"] W2 --> M2["Minutes each — budget survives a bad month, barely"] W3 --> M3["Hours — one event exceeds a year of budget"] M2 --> P2["Part 2: automate detection + recovery, contain failures"] M3 --> P2