fivenines
25/39

Theory Tutorial

Five nines where it counts — and nowhere else

Time
7m
Level
not specified
Artifacts
theory + practice
Progress0%
Lesson 3.1

Five nines where it counts — and nowhere else

How to assign each subsystem its own availability target so the composed system delivers five nines of trading without quintuple-nine everything.

The assembled four-nines venue has per-failure stories and a known durability gap. The next nine is too expensive to apply uniformly.

A plausible simplification is to label every service 99.999%, or spend five-nines engineering on all branches regardless of whether their outage stops trading. The catch is that the joint path that lets a user trade and see prices receives the five-nines budget; off-path work may lag or catch up as long as it cannot block or contradict that outcome.

Each nine costs roughly an order of magnitude more engineering than the last — spent uniformly, it bankrupts you. The escape comes from a definition: five nines of what? Define the availability contract as the user experiences it — "can I trade, and can I see prices?" — and suddenly most of the system doesn't need the last nine at all, because (1.1, 1.10) most of the system was already off that path:

SubsystemOutage meansTargetMechanism tier
Matching + sequencing + riskTrading stops99.999%Everything in Part 3
L1 market dataMarket goes blind99.999%Everything in Part 3
Order gateways (as a fleet)Some sessions reconnect99.999% fleet-wideMassive N+2 redundancy — already cheap (1.1)
L2/L3 depth, account APIsDegraded experience, trading continues99.99%Part 2 machinery as-is
Clearing, settlement, surveillanceBack-office lag, invisible for hours99.9%Part 1 machinery as-is — restart and catch up
Delayed data, analytics, web contentInconvenience99.5%Ordinary web engineering

Check the composition honestly (0.3's serial-chain math): a user's trading experience traverses gateway fleet → core → L1 feed. 0.99999 × 0.99999 × 0.99999 ≈ 99.997%… which fails the promise! Three five-nines components in series don't make five nines. The resolution: engineer the joint path, not the boxes. Failovers overlap (a zone loss takes gateway capacity and a core replica in the same seconds, not additively); the remaining independent-failure math then holds because each component's failures are sub-minute and rare. This is why five-nines systems are designed as one path with a shared budget — the table above allocates a single 5.3-minute budget across correlated components, roughly: 3 minutes core, 1 minute feeds, 1 minute edge, 0.3 reserve.

Everything below the five-nines line is a deliberate savings account: clearing at 99.9% with Part 1's restart-and-replay is a feature — complexity not built is reliability not lost. The discipline is writing the tier of each component down and defending the boundary when someone proposes "just making everything five nines to be safe."

Clearing is down for an hour. If matching waits, an off-path 99.9% subsystem destroys a five-nines trading promise. If obligations remain durable and replayable, trading availability and post-trade integrity can have different recovery times without conflict.

Define availability as the user's question: "can I trade and see prices?" Tier every subsystem; the last nine is only for outage-stops-trading components. Budget the joint path — serial math punishes box-by-box thinking.

Next step

See what actually stuck.

Take the practice scenarios now.