fivenines
20/39

Guided Problem

PFX Build 16: Change the Core Without Closing

Time
25m
Level
intermediate
Artifacts
not specified
Progress0%
Lesson 2.5

Change without downtime

How every tier upgrades while trading continues — and the compatibility discipline that makes rolling change possible at all.

A hot standby can take core ownership, while stateless edge fleets and replayable consumers already tolerate instance replacement.

The obvious shortcut is to upgrade both engines together and let new behavior activate when each process happens to restart. Yet one core version changes at a time, production history proves equivalent decisions before live ownership, adjacent versions interoperate, and behavior changes at a sequenced event.

Part 1 spent roughly half its downtime budget on planned maintenance. Four nines returns that budget by making deployment ride the redundancy we just built. Each tier has its own move:

  • Stateless edge — rolling replacement. Drain one gateway (stop new sessions; existing ones migrate per 2.4 — any gateway resumes any session), upgrade, verify, rejoin, next. The fleet absorbs the capacity dip invisibly.
  • Stateful core — upgrade by failover. You never upgrade a running engine; you upgrade the standby, let it prove itself consuming the live stream, then fail over to it (2.2). The old primary becomes the new standby on the old version — which is also your instant rollback: fail back.
  • Post-trade — pause and resume. Asynchronous consumers just stop, upgrade, and catch up from the journal. Their lag budget makes deployment trivial (1.10's payoff, again).
flowchart TB
    B["Release candidate passes replay test: yesterday's full production log, outputs diffed against production's"] --> U["Upgrade STANDBY to v2 — primary stays v1"]
    U --> W["Soak: v2 standby consumes live stream — divergence detector on (2.2)"]
    W --> D{"Outputs identical for N hours?"}
    D -- yes --> FO["Planned failover: v2 becomes primary, v1 standby remains as instant rollback"]
    D -- no --> AB["Abort — nothing user-visible happened"]
    FO --> OK{"Healthy after soak?"}
    OK -- yes --> FIN["Upgrade the v1 standby, done"]
    OK -- no --> RB["Fail back to v1 — seconds"]
    
Core upgrade pipeline. The deterministic log (1.7) gives an almost unfair superpower: the new version processes real production input — first replayed, then live-but-suppressed — and must produce bit-identical outputs before it's allowed to publish. Most bugs die in the diff, having never touched a user.

All of this hangs on one discipline: N and N−1 must interoperate, because every rolling or failover-based deploy has both versions live simultaneously. Protocol changes are additive (new optional fields, never repurposed ones); consumers ignore what they don't understand; message schemas are versioned; and state-machine changes that alter matching outcomes are gated on a sequenced activation event — "rule change R activates at event N" — so primary, standby, and any future replay all flip behavior at the same point in the stream. Without that last trick, a v2 standby would diverge from a v1 primary the moment their rules differed, and the divergence detector would (correctly) scream.

Replay diffing finds one different fill allocation. A red diff must never reach live trading, even if performance and health checks are green. The old primary remains authoritative until equivalence or an intentional sequenced rule change is established.

Edge rolls; core upgrades by failover; post-trade pauses and catches up. Replay-and-diff turns production history into the ultimate regression test. N/N−1 compatibility always; behavior changes activate via sequenced events.

Next step

See what actually stuck.

Take the practice scenarios now.