Operations
Disaster Recovery & Operations Runbook
Operational procedures for availability, failover, recovery, upgrade, and
reconciliation events. Everything here is a runnable drill — logs land in
evidence/ and the run is only green if every gate passes. These are ops
procedures; treat a green drill as evidence for that run, not a standing
RTO/RPO guarantee.
Topology
deployments/docker-compose.ha.yml — PostgreSQL primary + two streaming
replicas, HAProxy routing writes to whichever node is the read-write primary,
Prometheus + Alertmanager + exporter for metrics, RabbitMQ for the audit
outbox consumer.
The failover scripts assume this stack is up
(docker compose -f deployments/docker-compose.ha.yml up -d).
Failover drill — scripts/failover_payment_drill.sh
End-to-end economic proof of HA: baseline payment through the write proxy → kill the primary mid-flight → promote a replica → verify every acknowledged write survived on the new primary, replays don't double-post, and loan balances move exactly the expected delta. Then fences and rebuilds the demoted node.
Green run gate includes: refused write during the outage measured (the scripted gap is a drill metric, not a guaranteed RTO), all payment references present and unique on the promoted primary, allocations sum exactly to payment amounts, loan balances shift by the expected delta.
Latest verified run: evidence/failover-drill/failover_payment_drill_<ts>.log.
Recovery — scripts/recovery_drill.sh + cmd/recovery-verify
Backup → restore into a scratch database → recovery-verify walks the tables
and asserts the dump is internally consistent (row counts, referential
integrity, financial totals). Run after every backup-policy change and before
any upgrade. Artifacts under evidence/recovery-*/.
Tenant upgrade — scripts/upgrade_drill.sh + cmd/migrate-tenants
cmd/migrate-tenants is the canonical tenant migrator — applies all pending
db/migrations/tenant/* migrations per tenant schema in order. The upgrade
drill rehearse: snapshot → migrate → run probes → diff. Any tenant that
crosses 000300–000303 may surface FX repair flags — run
fx_flag_reconciliation_drill.sh on a staged tenant copy first, then review
real flags at Admin → FX Repair Flags after the production migrate.
FX reconciliation — scripts/fx_flag_reconciliation_drill.sh
See the FX Repair Flag Review Guide — the script is both the drill and the evidence procedure for real flags.
What an operator must check after any recovery event
audit_trailcontinuity — the hash chain must verify end-to-end; a gap or mismatch means events were lost between outbox and trail.- Reconciliation: subledger totals,
loan_balances, and journal nets must match the pre-event baseline for every tenant touched. - FX repair flag queue — any flags raised since the last check reviewed before period close.
- RabbitMQ consumers alive — outbox drain lag shows up as pending events in
audit_event_outbox.
Not covered here (explicit open items)
- Document custody / key recovery — tracked separately in the roadmap.
- Hosted-CI-gated release + production-scale soak beyond
evidence/soak-*. - Provider sandbox credentials and named-pilot sign-off — external.