Recommended Free Tools
A three-second failure detection followed by 13 seconds of continued traffic is not a Kubernetes-wide default. It is an incident-specific result produced by several independent timers: node heartbeats and health checks, the node controller’s grace period, taints and Pod tolerations, EndpointSlice updates, and the load balancer or proxy that consumes those updates.
Kubernetes’ published defaults describe a different sequence. The node controller checks state every five seconds, normally waits five minutes after marking a node Unknown before submitting the first eviction request, and automatically gives Pods 300-second tolerations for node.kubernetes.io/not-ready and node.kubernetes.io/unreachable unless those tolerations are changed.
What the 3-second and 13-second intervals actually mean
The two intervals measure different events. “Detected in three seconds” may be the time taken by a custom node health check, a cloud load balancer, a monitoring system, or a particular controller configuration to notice a failure. It is not established as the upstream Kubernetes node-failure detector’s default.
“Keeps receiving traffic for 13 seconds” is also not determined solely by node detection. Even after Kubernetes changes a Pod’s state, the control plane must update EndpointSlices and the selected data plane must program or refresh its backend list. Existing connections may continue while a proxy drains them. A network partition can extend the problem further because the control plane may be unable to make the kubelet stop a process on the isolated node.
#1 Best Overall
How Kubernetes detects an unhealthy node
Heartbeats and node conditions
Kubernetes nodes send heartbeats through their status updates and leases. The node controller uses those signals to assess availability and change the Node condition when communication stops. A heartbeat loss is therefore a control-plane observation, not proof that every process on the host has stopped.
The five-second state check is not a three-second failure timeout
The Kubernetes Nodes documentation lists a five-second default interval for checking node state. That interval is only one clock in the sequence. The controller also applies grace periods and health-signal requirements, and managed Kubernetes services can change controller settings. The node lifecycle controller source comments indicate that the monitoring grace period must allow multiple health-signal intervals; the appropriate value is release-dependent.
Unknown, taints and eviction are later stages
After a node is considered unavailable, Kubernetes can set conditions and add taints such as node.kubernetes.io/not-ready or node.kubernetes.io/unreachable. The default delay before the first eviction request after a node is marked Unknown is five minutes, not three seconds. The documented default eviction rate is 0.1 nodes per second in most cases, so a large failure can take longer to process.
Rank #2
Why Pods may continue running on an unreachable host
Kubernetes automatically adds 300-second tolerations for the not-ready and unreachable taints unless a Pod specification or its controller changes them. A Pod that tolerates the taint can remain bound to the node during that period.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minutePC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11A partition is especially important. If the control plane cannot communicate with the kubelet, an API-level deletion or eviction does not guarantee that the container process on the isolated machine has exited. The process may continue serving requests even though Kubernetes shows the node as unhealthy. Treat “Pod deleted” and “application stopped” as separate observations until the host or runtime is independently confirmed.
Why traffic can continue after endpoint changes
Terminating endpoints should be removed from normal selection
For regular Service traffic, an EndpointSlice endpoint that is terminating has ready=false. Load balancers should not select such an endpoint for new traffic. This behavior depends on the consumer actually receiving and applying the EndpointSlice update.
Draining existing connections
The EndpointSlice serving condition can be used to represent whether a terminating endpoint can still serve existing connections. A proxy or load balancer may therefore stop new requests quickly while allowing established connections to drain, making “traffic continued” appear longer than endpoint selection for new requests.
Data-plane and external load-balancer lag
After the API object changes, kube-proxy, a service mesh, an ingress controller, or an external load balancer must program its own state. Refresh intervals, watch delivery, retries, connection reuse and provider-specific health checks determine when the backend is actually removed. The 13-second observation cannot be attributed to Kubernetes alone without timestamps from that data plane.
Free tools Windows power users keep installed
One-click scans. No signup required.
Published defaults versus the incident’s measurements
| Stage | Published value or behavior | What it does not prove |
|---|---|---|
| Node-controller state check | Five seconds by default | That a node is declared dead exactly three seconds after failure |
First eviction request after Unknown |
Five minutes by default | That a Pod leaves traffic backends within five minutes, or that eviction stopped the process |
| Automatic not-ready/unreachable toleration | 300 seconds by default | That every Pod remains for 300 seconds; Pod or controller tolerations can differ |
| Node eviction rate | 0.1 nodes per second in most cases | That a multi-node failure is handled at a fixed per-Pod speed |
| Terminating EndpointSlice endpoint | ready=false for normal traffic selection |
That every proxy or external load balancer applies the change immediately |
How to determine what happened in your cluster
Build one timeline using UTC timestamps. Do not infer the cause from a single event or from the time a monitoring alert was delivered.
Rank #4
- Identify the implementation. Record the Kubernetes version, distribution, control-plane settings, CNI, kube-proxy mode, service mesh, ingress and external load balancer.
- Capture node signals. Compare the last node lease or status update with Node condition transitions and taint timestamps.
- Check Pod behavior. Inspect Pod start, deletion and termination timestamps, the owning controller, and the Pod’s tolerations for not-ready and unreachable taints.
- Inspect EndpointSlices. Record when each endpoint changed
ready,servingandterminatingstate. - Inspect the data plane. Find when kube-proxy, the mesh, ingress or provider load balancer received and programmed the backend change.
- Verify the host separately. If a partition is possible, check the node runtime, process table, network path and application logs directly; an API deletion alone is not proof of shutdown.
Useful starting commands are:
kubectl get node <node> -o yamlto inspect conditions, leases-related metadata and taints.kubectl describe node <node>to review conditions and events.kubectl get pod -A -o wideto locate Pods still assigned to the node.kubectl get pod <pod> -o yamlto inspect deletion timestamps, termination state and tolerations.kubectl get endpointslice -A -o yamlto compare endpoint readiness and serving conditions.kubectl get events -A --sort-by=.lastTimestampto correlate controller actions, while retaining the original event timestamps because event retention is limited.
Common explanations for a short detection and longer traffic window
A custom health check detected the host first
A three-second alert can come from a node agent, a cloud health check or an application-side monitor that runs more frequently than the Kubernetes node controller. In that case, the alert and Kubernetes’ own condition transition are measuring different systems.
Endpoint propagation took longer than failure detection
The node can be declared unhealthy quickly while EndpointSlice watches, proxy programming or an external provider take additional seconds. The relevant evidence is the timestamp of each state transition, not the alert timestamp alone.
Connections were draining
If new backend selection stopped but existing keep-alive or streaming connections remained, the observed 13 seconds may be connection-drain time rather than continued selection of a dead endpoint.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →The node was partitioned, not powered off
A partition can leave the application process alive and reachable to some clients while the control plane considers the node unreachable. This explains traffic continuing after Kubernetes actions that the isolated kubelet never received.
Operational implications
- Set Pod tolerations deliberately for applications that must fail over quickly; the 300-second automatic toleration is not a universal availability target.
- Make readiness and termination behavior cooperate with your proxy’s draining model, especially for long-lived connections.
- Monitor node heartbeats, EndpointSlice changes and actual load-balancer backends as separate signals.
- Test node failure and network-partition scenarios independently. A powered-off node, a blocked kubelet connection and a failed application process produce different timelines.
- When documenting an incident, label every duration with its source: alert, Node condition, taint, eviction request, endpoint update or backend removal.
The Bottom Line
A dead node being noticed in three seconds and still receiving traffic for 13 seconds is plausible as a system-specific incident, but neither number is a general Kubernetes default. The five-second state-check interval, five-minute default eviction delay and 300-second automatic tolerations describe separate control-plane stages; endpoint propagation, connection draining and network partitions determine when traffic truly stops.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

