The event backbone: outbox, idempotency, replay
From 1.14: sync call chains cascade failures, and 0.3 says serial chains multiply unavailability. From 1.7: trip completion triggers payments, comms, ratings, history, analytics — five downstream effects. Today those are calls; if comms is down, does completion fail? It must not.
The simplest version would commit trip state and publish its event as two ordinary operations because both are usually fast. The better question is what survives when that assumption fails. A committed business fact cannot lose its event or create a phantom event, and redelivery produces one consumer effect.
The architectural move of Part 2: a trip transition is a fact, not a procedure call. The Trip service commits its state change and, in the same database transaction, appends the event to an outbox table (this atomicity is the whole trick — no distributed transaction, no dual-write race). A relay tails the outbox and publishes to the event log, partitioned by trip id so each trip's events stay ordered. Downstream services — payments, comms, history, ratings, analytics — become independent consumers with independent offsets. Trip completion now succeeds even if every consumer is down; consumers catch up when they recover, and the log retains days of history so recovery is replay, not archaeology.
The price of this durability is at-least-once delivery: consumers will see duplicates (relay retries, consumer restarts, rebalances). The idempotency discipline from 1.8 and 1.9 now becomes a system-wide law: every consumer keys its effect on (event id, consumer) and makes redelivery a no-op. Exactly-once effects over at-least-once delivery — through idempotency, not through belief in exactly-once delivery. Poison events that repeatedly crash a consumer go to a dead-letter queue with an alert, never blocking the partition behind them.
sequenceDiagram
autonumber
participant T as Trip service
participant DB as Trip DB (shard)
participant REL as Outbox relay
participant K as Event log
participant PAY as Payments consumer
participant CM as Comms consumer
T->>DB: one transaction: state = Completed + outbox row
DB-->>T: committed — trip is done, full stop
REL->>DB: poll outbox
REL->>K: publish TripCompleted (partition by trip id)
REL->>DB: mark outbox row sent
K-->>PAY: TripCompleted
PAY->>PAY: seen event id? no → capture fare (1.8)
K-->>CM: TripCompleted
CM--xCM: comms is down!
Note over CM,K: offset not advanced — event waits in the log
CM->>K: recovered — resume from offset
K-->>CM: TripCompleted (redelivered)
CM->>CM: dedupe on event id → send receipt once
flowchart LR
T["Trip service"] --> OB[("outbox table — same tx as state")]
OB --> REL["Relay"]
REL --> K[["Event log — trip-events, partitioned by trip id, 7 d retention"]]
K --> PAY["Payments"]
K --> CM["Comms"]
K --> HIST["History builder"]
K --> RATE["Ratings"]
K --> AN["Analytics 2.10"]
PAY --> DLQ[["Dead-letter queue + alert"]]
CM --> DLQ
Consumer lag is now a first-class SLI. The system "works" while receipts drift 40 minutes behind — invisible on error dashboards, very visible to users. Alert on lag per consumer group (symptom: "receipts delayed", not just "lag > N"), and size consumers to drain a worst-case backlog quickly: recovery speed is part of availability.