Operations

Disaster Recovery & Operations Runbook

Operational procedures for availability, failover, recovery, upgrade, and reconciliation events. Everything here is a runnable drill — logs land in evidence/ and the run is only green if every gate passes. These are ops procedures; treat a green drill as evidence for that run, not a standing RTO/RPO guarantee.

Topology

deployments/docker-compose.ha.yml — PostgreSQL primary + two streaming replicas, HAProxy routing writes to whichever node is the read-write primary, Prometheus + Alertmanager + exporter for metrics, RabbitMQ for the audit outbox consumer.

The failover scripts assume this stack is up (docker compose -f deployments/docker-compose.ha.yml up -d).

Failover drill — scripts/failover_payment_drill.sh

End-to-end economic proof of HA: baseline payment through the write proxy → kill the primary mid-flight → promote a replica → verify every acknowledged write survived on the new primary, replays don't double-post, and loan balances move exactly the expected delta. Then fences and rebuilds the demoted node.

Green run gate includes: refused write during the outage measured (the scripted gap is a drill metric, not a guaranteed RTO), all payment references present and unique on the promoted primary, allocations sum exactly to payment amounts, loan balances shift by the expected delta.

Latest verified run: evidence/failover-drill/failover_payment_drill_<ts>.log.

Recovery — scripts/recovery_drill.sh + cmd/recovery-verify

Backup → restore into a scratch database → recovery-verify walks the tables and asserts the dump is internally consistent (row counts, referential integrity, financial totals). Run after every backup-policy change and before any upgrade. Artifacts under evidence/recovery-*/.

Tenant upgrade — scripts/upgrade_drill.sh + cmd/migrate-tenants

cmd/migrate-tenants is the canonical tenant migrator — applies all pending db/migrations/tenant/* migrations per tenant schema in order. The upgrade drill rehearse: snapshot → migrate → run probes → diff. Any tenant that crosses 000300–000303 may surface FX repair flags — run fx_flag_reconciliation_drill.sh on a staged tenant copy first, then review real flags at Admin → FX Repair Flags after the production migrate.

FX reconciliation — scripts/fx_flag_reconciliation_drill.sh

See the FX Repair Flag Review Guide — the script is both the drill and the evidence procedure for real flags.

What an operator must check after any recovery event

  1. audit_trail continuity — the hash chain must verify end-to-end; a gap or mismatch means events were lost between outbox and trail.
  2. Reconciliation: subledger totals, loan_balances, and journal nets must match the pre-event baseline for every tenant touched.
  3. FX repair flag queue — any flags raised since the last check reviewed before period close.
  4. RabbitMQ consumers alive — outbox drain lag shows up as pending events in audit_event_outbox.

Not covered here (explicit open items)

  • Document custody / key recovery — tracked separately in the roadmap.
  • Hosted-CI-gated release + production-scale soak beyond evidence/soak-*.
  • Provider sandbox credentials and named-pilot sign-off — external.