What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
Redis Sentinel failover has no single documented end-to-end duration. Its elapsed time depends on how quickly Sentinels detect a failure, whether they can agree and authorize a failover, how quickly a replica is promoted and the topology is reconfigured, and when your clients reconnect and resume work. Redis documentation explains these stages but does not give a universal Sentinel failover SLA.
No measured duration or test method is available here, so there is no defensible number to report as a result. To measure a real deployment, define the starting event and recovery endpoint, then report the Redis version, Sentinel configuration, topology, failure trigger, and client behavior.
What counts as a Sentinel failover’s duration?
“Failover time” can describe different intervals. A useful report should state whether its stopwatch starts when a server becomes unreachable, when a Sentinel marks it down, or when an application first observes an error. It should also identify whether the endpoint is replica promotion, topology reconfiguration, successful client reconnection, or restored application requests.
Free tools Windows power users keep installed
One-click scans. No signup required.
Those endpoints are not interchangeable. Sentinel may have promoted a replica while a client is still using a stale connection or has not discovered the new primary. Redis’s Sentinel documentation describes the monitoring, agreement, and failover mechanisms; it does not establish one end-to-end duration for all deployments.
#1 Best Overall
How Sentinel timing affects recovery
Failure detection: SDOWN and ODOWN
A Sentinel marks a monitored instance subjectively down (SDOWN) after it has not received a valid PING response for the configured down-after-milliseconds interval. This is a local judgment, not yet a cluster-wide decision. Sentinels mark an instance objectively down (ODOWN) when enough other Sentinels agree to meet the configured quorum.
Detection can therefore take longer than the configured threshold if the relevant Sentinel processes cannot communicate or obtain the required agreement. The threshold is one part of the timing, not a complete failover timer.
Rank #2
Authorization: quorum is not the whole vote
Agreement that the primary is ODOWN and authorization to perform a failover are distinct requirements. A failover also needs authorization from a majority of Sentinel processes. A deployment can detect an outage yet be unable to proceed if Sentinel communication or voting is impaired.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minuteReplica selection, promotion, and reconfiguration
After authorization, Sentinel chooses a suitable replica, promotes it, and reconfigures the remaining topology. Selection considers factors such as how long a replica has been disconnected, its configured priority, replication offset, and run ID. A replica’s state and lag can affect which candidate is suitable and how quickly the new topology becomes usable.
Rank #3
The parallel-syncs setting controls how many replicas Sentinel reconfigures to follow the promoted replica at the same time. It affects replica availability during synchronization; it is not a direct guarantee of how long the primary promotion or client recovery will take. Redis documents these behaviors in its Sentinel guide and the sentinel.conf source file.
Why failover-timeout is not a duration guarantee
failover-timeout is a Sentinel control parameter, not a promise that the entire failover will finish within that many milliseconds. Redis documents multiple uses for it, including retry timing after an earlier attempt and waiting periods within the failover process. Its value should not be reported as an observed or guaranteed recovery time.
Do not confuse Sentinel failover with the FAILOVER command
Redis’s FAILOVER command documentation says: “Failovers typically happen in less than a second, but could take longer if there is a large amount of write traffic or the replica is already behind in consuming the replication stream.” That statement concerns the coordinated FAILOVER command context. It is not a measured or promised duration for Sentinel’s full failure-detection-to-client-recovery path.
Defaults are not measured failover times
The inspected Redis unstable-branch source declares a 1-second Sentinel ping period, a 30-second default down-after-milliseconds threshold, and a 180-second default failover-timeout. These are implementation defaults in that branch, not universal settings across Redis versions or deployments, and none is an observed end-to-end failover duration. The values are visible in sentinel.c; verify the version and effective configuration of the Redis installation you are measuring.
How to report a real failover measurement
A useful measurement describes both the environment and the boundary of the stopwatch. Include:
- Version and configuration: Redis version and effective Sentinel timing settings, including
down-after-milliseconds, quorum,failover-timeout, andparallel-syncs. - Topology: Number of Sentinel processes, monitored primary and replicas, and which replica was promoted.
- Failure trigger: Exactly what was interrupted and how, along with relevant network conditions.
- Timing boundaries: The event that starts the clock and the specific endpoint that stops it, such as first successful application request to the new primary.
- Client recovery: Client library and reconnection or discovery behavior, because Sentinel’s topology change alone does not show when the application can use it.
- Run-to-run variation: Results across repeated runs, with the same trigger and timing boundaries, rather than a single value presented as generally representative.
When comparing deployments, keep the failure trigger and measurement boundaries consistent. Differences in Sentinel count, quorum, network reachability, replica lag and client recovery can otherwise make two reported “failover times” measure different things.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

