Replication & Failover Recovery
Replication broke, a failover went wrong, or two nodes both think they are primary. We restore a single source of truth and rebuild redundancy.
Signs you are here
- Replicas are far behind or stopped applying changes
- After a failover, writes went to two places
- The standby will not promote or rejoin
Until someone has taken a copy, avoid restarts, repair commands, and restores over the original.
How we approach it
- Find the truth Determine which node holds the correct, most complete data.
- Reconcile Recover writes that landed on the wrong node.
- Rebuild redundancy Re-seed replicas and verify they stay in sync.
- Test failover Run a controlled failover so you know it works before the next incident.
Questions
Can you work with managed cloud databases?
Yes — AWS RDS and Aurora, Google Cloud SQL and Azure SQL, within what each provider exposes.
Engines we work with
- PostgreSQL
- MySQL
- MariaDB
- SQL Server
- AWS RDS & Aurora
- Google Cloud SQL
- Azure SQL