Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

A safe failover test begins by confirming which server is leader and which server is intended to take over. It ends only after you verify that the former leader cannot accept writes and has rejoined safely. The title does not identify a particular incident or platform; the steps below use Patroni and PostgreSQL as a documented example, not as an attribution of what happened in any specific test.

Why can a failover test stop the wrong node?

“The wrong node” can mean the current leader was stopped instead of the intended candidate, the wrong replica was promoted, or an old primary came back and began accepting writes after a new primary had been promoted. These are different failure modes, so first define what the exercise is meant to test: a planned switchover, primary failure, loss of access to the distributed configuration store (DCS), a network partition, or disaster recovery at a separate site.

With Patroni, the cluster’s leader lock coordinates which PostgreSQL instance may be primary. Patroni attempts to stop PostgreSQL if it cannot renew that lock, but a separate process manager that independently restarts PostgreSQL can undermine that protection. The Patroni FAQ states: “Only Patroni should be able to start, stop and promote Postgres instances in the cluster.” Read the Patroni FAQ.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What to verify before initiating failover

Confirm the leader and candidate

Capture the cluster status before making any change. Record each member’s name, role, and state, then confirm the leader identity and that the proposed candidate is both the intended machine and healthy enough to promote. Patroni’s cluster API provides member roles and states; its manual failover request requires a named candidate. Consult the Patroni REST API documentation for the endpoints and parameters supported by the deployed release.

Define the failure and data-loss objective

Choose the event being simulated and set an acceptable recovery point objective (RPO): how much recent data, if any, the exercise can tolerate losing. Patroni’s manual failover API warns that failover can cause data loss. Patroni describes asynchronous replication as the default and provides a configurable maximum-lag threshold, but a threshold is not a guarantee that every recent write has reached a replica. Decide in advance how to identify missing or divergent writes after promotion. See the Patroni project README.

Check who can control PostgreSQL

In a Patroni-managed cluster, ensure that no independent service-manager restart policy can start PostgreSQL on a former primary. The cluster manager must control database startup, shutdown, and promotion consistently; otherwise, a stopped old primary may return as a second writable server.

Validate fencing and its failure behavior

Fencing is the mechanism that makes an old primary unable to write, rather than merely assuming it has stopped. Patroni supports a pre_promote hook that runs after the leader lock is acquired and before PostgreSQL is promoted. If the hook exits unsuccessfully, Patroni blocks promotion and removes the leader key. Test the hook’s actual action and failure path before a live exercise; a hook that reports success without isolating the old primary does not provide effective fencing. Review Patroni’s replication and promotion documentation.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How to run a controlled Patroni failover exercise

  1. Record the baseline. Use the supported cluster status interface to save member names, roles, states, and the current leader. Check that the named candidate is healthy and that its replication position is acceptable under the agreed RPO.
  2. Confirm the operation. For a planned transfer, use the documented switchover operation. For a manual or emergency promotion, follow the manual failover API’s requirements, including explicitly naming the candidate. Do not treat these as interchangeable procedures: a manual failover request can be used while a leader exists and carries a data-loss warning. Use the deployed Patroni version’s REST API documentation.
  3. Apply only the intended fault. Carry out the chosen scenario—such as a planned switchover or simulated primary loss—without adding unrelated failures. In a network-partition or DCS-loss exercise, observe which members can still reach the coordination system and whether the old leader is actually prevented from writing.
  4. Watch status and client-facing health. Track cluster membership, leader identity, lock validity, replica catch-up, and the point at which application connections recover. Patroni health and readiness endpoints can help distinguish primary status from replica readiness; use the endpoint behavior documented for your release rather than assuming that a process being alive means it is ready to serve the intended traffic.
  5. Verify the old primary is isolated. Before declaring success, confirm that the former leader cannot accept writes, that only one server is writable, and that replicas follow the new leader. Check application data against the RPO and look for missing or divergent writes.
  6. Restore redundancy deliberately. Rejoin the former primary through the documented recovery path and verify it follows the new leader. Promotion is not the end of the exercise: Patroni notes that redundancy is reduced until the failed member returns.

What changes in a two-site disaster-recovery test?

A second site using asynchronous replication cannot necessarily determine whether the source site is down or merely unreachable. In the documented two-site standby arrangement, automatic promotion is not possible because the standby site cannot know the source site’s state. Confirm that the source is down and fence it—often described as STONITH—before promoting the standby. Promoting while the source cluster remains active can create split brain; the Patroni guide puts it plainly: “If the source cluster is still up and running and you promote the standby cluster you create a split-brain.” Read Patroni’s multi-datacenter guide.

After the source site is recovered, reconcile the topology before restoring normal service. Do not assume that restarting its former primary makes it safe to rejoin; establish which cluster is authoritative and use the supported recovery procedure for the deployment.

When is a watchdog useful?

A watchdog is an additional protection for cases where Patroni cannot reliably stop PostgreSQL—for example, if the agent crashes, is killed, runs too slowly, or the virtual machine is paused or heavily loaded. Depending on the setup, watchdog expiry can reset the host if the agent fails to maintain its safety conditions. It supplements rather than replaces correct leader coordination and fencing.

Watchdog timing is tied to the DCS leader-lock time-to-live (TTL). Patroni’s documentation shows example defaults, including a 30-second TTL and a five-second safety margin, but these are configuration examples, not universal settings or performance guarantees. Check the actual loop_wait, retry_timeout, and ttl values together, and validate their timing on the deployed system rather than copying defaults. See Patroni’s watchdog documentation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What makes the test pass?

  • The intended candidate—not an unverified member—became the leader.
  • At no point were two PostgreSQL instances able to accept writes.
  • The old primary remained fenced or stopped, including under the tested failure conditions.
  • Client-facing health matched the role and readiness expected for each member.
  • Data was checked against the stated RPO, and any missing or divergent writes were understood.
  • The former primary rejoined safely and redundancy was restored.

Match every command, endpoint, and recovery action to the Patroni release and topology actually deployed. The project documentation cited here tracks a mutable branch, so operational details may change between releases.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.