When a Kubernetes node stops reporting, the control plane eventually marks it Ready=Unknown and adds an unreachable taint. Most ordinary pods tolerate that taint for 300 seconds by default before they become eligible for eviction. A workload controller may then create a replacement pod, but it cannot move the original pod to another node. If the node is merely cut off from the control plane, its old process may still be running even after the pod is deleted from the API.
What happens, step by step?
- The node stops sending heartbeats. Kubernetes monitors nodes using status updates and Lease objects. If neither is renewed, the control plane treats the node as potentially unavailable. See the Kubernetes Nodes documentation.
- The node is marked unknown after a grace period. Once the configured
node-monitor-grace-periodexpires, the node controller sets the node’sReadycondition toUnknown. Kubernetes documents a default of 50 seconds, but administrators can configure a different value. This is a detection threshold, not an end-to-end recovery guarantee. See Node Status. - The control plane applies an unreachable taint. The
node.kubernetes.io/unreachabletaint hasNoExecutebehavior by default. This can prevent new pods from being placed there and, for existing pods without a matching toleration, make them eligible for eviction. See Taints and Tolerations. - Each pod’s toleration determines its eviction timing. Most ordinary pods receive a default 300-second toleration for unreachable and not-ready taints. The eviction clock starts when the taint is applied, not when the original failure began. Explicit tolerations can shorten or extend the delay, or omit a finite limit.
- Eviction and replacement are separate actions. Taint-based eviction can delete the pod object through the API. An owning controller may then create a replacement to restore the desired replica count, subject to scheduling, capacity, storage, and workload constraints.
The 50-second node-monitor grace period and 300-second default pod toleration are distinct intervals. They should not be added together as a universal recovery promise: settings and controller behavior vary, and replacement scheduling may take longer or fail.
How tolerations and workload type change the result
| Pod or setting | Effect when the unreachable taint is applied |
|---|---|
| Ordinary pod with default toleration | Normally remains bound during the default 300-second toleration, then becomes eligible for eviction. |
Pod with no matching NoExecute toleration |
Eligible for eviction without a configured waiting period. |
Pod with finite tolerationSeconds |
Eligible for eviction after the configured duration. |
Pod with a matching toleration and no tolerationSeconds |
Can remain bound indefinitely while the taint remains. |
| DaemonSet pod | DaemonSet pods tolerate unreachable and not-ready taints indefinitely, so these taints do not evict them. |
Check the pod’s actual tolerations rather than assuming defaults: a pod or controller can specify tolerations that change or replace the ordinary behavior. Since Kubernetes 1.29, taint-based eviction is handled by a separate taint-eviction-controller; cluster configuration can disable it, so the deployed release and configuration matter.
Does Kubernetes restart the pod on another node?
No. Kubernetes does not transfer a pod’s binding to a different node. A pod is identified by its UID, and a replacement is a new pod with a different UID, even if it has the same name and specification. The Pod Lifecycle documentation explains that a pod is replaced rather than rescheduled.
#1 Best Overall
A Deployment, ReplicaSet, StatefulSet, Job, or another owning controller may create a new pod if its desired state requires one. The scheduler chooses a node based on available capacity, affinity and topology rules, and other constraints. Stateful workloads may also wait for storage or require additional recovery steps; a replacement is not guaranteed to start immediately or on the same node.
Why an unreachable node is different from a powered-off node
Missing heartbeats do not tell Kubernetes whether a machine is shut down or merely isolated from the control plane. During a network partition, the API server may record a pod deletion but be unable to deliver that request to the kubelet on the isolated node. The old process may therefore continue running while a replacement starts elsewhere. Kubernetes documents this possibility in Taints and Tolerations.
This matters most for stateful workloads. If both the old and replacement processes can write to the same data or act as the active leader, the result may be conflicting writes or corruption. Consider application-level leadership or leases, node fencing, and the storage system’s ownership guarantees before forcing recovery. A pod’s disappearance from the API is not proof that its process has stopped.
What operators should check first
- Inspect node conditions and taints with
kubectl describe node <node-name>. The Nodes documentation describes this command for viewing node conditions. - Check pod placement and status with
kubectl get pods -o wide, then inspect the affected pod’s tolerations and owner references. - Check whether a replacement exists and, if it is Pending, investigate capacity, affinity, topology, and volume constraints.
- For stateful workloads, establish whether the old machine is actually shut down and confirm the fencing and storage recovery policy before taking actions that could permit a second writer.
A PodDisruptionBudget is generally used for voluntary disruptions through the eviction API. Hardware failure and a network partition are involuntary disruptions, so a PDB should not be treated as protection against node-failure eviction. See Pod Disruptions.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →Rank #3
When force deletion or volume detach is appropriate
If a node has genuinely shut down non-gracefully and a volume remains attached, Kubernetes documents an node.kubernetes.io/out-of-service taint workflow for force-deleting pods and detaching volumes. This is an administrator procedure, not a safe shortcut for an ambiguous outage. Verify that the node is shut down and not about to restart before applying the taint; remove it after the node has recovered and migrated pods have been checked. Kubernetes warns that force-detaching a volume while the old workload might still be active can violate storage ordering expectations and risk data corruption. See Non-Graceful Node Shutdown.
The same documentation describes optional, configuration-dependent force-detach behavior after a six-minute deletion timeout. That timing is not universal and does not remove the need to verify that the old workload cannot still write.
Rank #4
What determines the actual recovery time?
- The configured node-monitor grace period and the time until the failure is detected.
- The pod’s matching tolerations and whether the taint-eviction controller is enabled.
- The owning workload controller’s behavior and desired state.
- Available node capacity and scheduler constraints.
- Whether storage can attach safely and whether the workload requires fencing or application-level recovery.
- Cloud-provider and storage-driver behavior in the specific cluster.
Kubernetes documents default mechanisms and settings, not a universal time from node failure to a healthy replacement.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

