Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To explain a Kubernetes production outage, identify what its controller was meant to keep running, compare that desired state with what Pods, nodes, storage, and the application actually did, and trace where recovery stopped. A Deployment, DaemonSet, or StatefulSet can shape an incident’s scope and recovery, but the controller type alone does not establish the cause. Without a cluster version, service, timeline, or incident evidence, no specific outage or root cause can be attributed here; the guide below provides a controller-aware way to investigate one.

Deployment vs. StatefulSet vs. DaemonSet

These controllers solve different placement and identity problems. Choose based on what a workload needs from its Pods, not on a general ranking of which controller is safer. Kubernetes describes the broader workload options in its Workloads documentation.

Decision point Deployment DaemonSet StatefulSet
Pod relationship Generally interchangeable replicas managed through ReplicaSets. A local instance on each matching node, or on a selected subset of nodes. Pods have stable, unique ordinal identities; their identities and storage associations are not interchangeable in the same way.
Placement The scheduler places replica Pods subject to scheduling rules. Node labels, selectors, and scheduling eligibility determine which nodes should run a Pod. The scheduler places Pods while the controller preserves identity and, depending on policy, ordering semantics.
Typical fit Stateless frontends and APIs, or other workloads whose replicas can serve equivalent work. Node-local facilities such as network plugins, logging agents, or storage agents. Workloads that need stable identity, persistent-claim association, or ordered behavior.
Update and recovery focus Inspect the ReplicaSet rollout, progress, and retained revisions. Inspect update status and which eligible nodes have received the new Pod. Inspect Pod ordinals and readiness; an unready Pod can block later progress under ordered updates.
Persistence semantics Deployment semantics do not themselves provide persistent storage. DaemonSet semantics do not themselves provide persistent storage. With volumeClaimTemplates, the controller can maintain stable identity-to-claim associations; storage availability and data safety still require separate handling.

These behaviors are described in the Kubernetes documentation for DaemonSets, StatefulSets, and the Apps API.

When should I use a Deployment?

Use a Deployment when replicas are generally interchangeable and you want declarative scaling and a managed rollout without assigning each Pod to a particular host. Kubernetes describes Deployments as managing Pods through ReplicaSets. A Deployment is a poor fit when the application depends on a specific stable Pod identity or a node-local copy on every matching host.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

When should I use a DaemonSet?

Use a DaemonSet when a copy must run on every node matching the workload’s selection and scheduling rules, or on a defined subset. It is commonly used for node-local networking, logging, and storage functions. Do not equate “one per matching node” with “one per every node”: label changes, taints, tolerations, resource pressure, and other scheduling constraints can change which nodes are actually eligible.

When should I use a StatefulSet?

Use a StatefulSet when Pods need stable identities, stable storage associations, or ordered deployment and scaling behavior. Its identity helps the system associate a replacement Pod with the same ordinal and, when configured, its claim. It does not make the application highly available, ensure a volume is attachable, or guarantee that the application can recover its data.

What a controller can—and cannot—tell you about an outage

A controller reconciles declared desired state. A healthy-looking controller status is not proof that the service is healthy: replica counts, Pod readiness, endpoint membership, application behavior, and storage recovery are distinct observations. Kubernetes can create a replacement Pod to maintain a requested replica count when a Pod fails, but it cannot repair an application defect or every storage failure. See the official explanation of Kubernetes self-healing.

For an incident report, establish a timeline from actual evidence: what changed, when symptoms began, what the controller reported, which Pods or nodes were affected, what users observed, and how service returned. The controller may have exposed or amplified a failure, or may simply have been responding to it. Do not assign causation or quantify affected replicas, nodes, shards, or recovery time without incident records.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How to investigate the failure in controller order

  1. Identify the owner chain

    Find the affected Pod’s controlling resource before editing anything: a Deployment usually owns a ReplicaSet that owns Pods; a DaemonSet or StatefulSet manages its Pods directly. Check selectors and ownership references. Overlapping selectors can make ownership and observed behavior confusing, so avoid changing resources until you know which controller is responsible.

  2. Compare desired state with observed state

    Record desired, current, ready, and available replicas where applicable, plus updated-replica counts and controller conditions. Preserve kubectl describe output and Events while the failure is occurring; they can capture scheduling, image, probe, and mount problems that are no longer visible after recovery. For a Deployment, check rollout progress with:

    kubectl rollout status deployment/<name> -n <namespace>

    The Deployment progress deadline defaults to 600 seconds in the current Kubernetes documentation retrieved in 2026. Exceeding it sets the Deployment’s Progressing condition to false; that condition is a signal to investigate Events and Pod startup, not a root-cause explanation. The same documentation gives default Deployment RollingUpdate values of maxUnavailable: 25% and maxSurge: 25%; percentage rounding is down for maxUnavailable and up for maxSurge. Check the relevant Deployment rolling-update guidance and the actual workload configuration rather than assuming defaults were left unchanged.

  3. Check the controller’s scope and placement

    For a DaemonSet, enumerate the nodes that match its selector and check eligibility: node labels, taints and tolerations, and resource pressure can explain why the expected node-local Pod is absent or unscheduled. For a StatefulSet, line up each ordinal with its Pod, PersistentVolumeClaim, and stable DNS identity. Compare the affected object and its neighbors instead of treating a controller-wide count as a complete service diagnosis.

    Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  4. Separate rollout failure from runtime failure

    A new Pod can be created and still fail readiness, crash, or serve incorrect results. Correlate image and configuration changes with probe results, Events, container logs, Service endpoints, and application-level health. A ready Pod is useful evidence, but it does not by itself prove the user-facing request path is working.

  5. Trace storage and dependencies

    For stateful workloads, check claim and volume binding, attachment, mount errors, the StorageClass and provisioner, and the application’s own data-recovery behavior. Stable StatefulSet identity can make it clearer which claim belongs to a replacement Pod; it does not make the volume available or the data safe. Also examine external dependencies the application needs to become ready or serve requests.

  6. Quantify impact from incident records

    Use service telemetry, controller and Pod history, node state, and application records to establish the blast radius and duration. A controller’s desired count is not a measure of the number of users or requests affected. No attributable population statistic or named study in the cited source set establishes an outage rate, mean recovery time, or failure percentage for these controllers.

How rollout and rollback behavior changes by controller

Deployment: ReplicaSet-based rollout

A Deployment’s rolling update uses maxUnavailable and maxSurge to govern unavailable and additional Pods during a rollout. The current documented defaults are stated above; verify the configured values and available capacity for the incident workload. The current documentation also gives a default retained history of 10 old ReplicaSets. Setting revisionHistoryLimit: 0 disables rollback, so retained history should be checked before relying on revision-based recovery.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To roll back to the previous retained Deployment revision, first verify that doing so is appropriate for the incident, then run:

  1. kubectl rollout history deployment/<name> -n <namespace>
  2. kubectl rollout undo deployment/<name> -n <namespace>
  3. kubectl rollout status deployment/<name> -n <namespace>

Use --to-revision=<revision> with rollout undo only after identifying the intended retained revision. A rollback restores a prior Pod template; it does not necessarily reverse external side effects, data migrations, or configuration changes outside that template. Follow the official Deployment update and rollback instructions.

DaemonSet: node-scoped rollout

A DaemonSet update propagates across eligible nodes. During an incident, determine how broadly the new version was applied and whether the matching node set changed; reverting a DaemonSet can affect node-local functionality across that same scope. Verify the DaemonSet’s update strategy, status, node labels, and actual Pod versions against the cluster’s configuration. The DaemonSet documentation covers node matching and update behavior.

StatefulSet: ordered progress and recovery

With ordered update behavior, rollout proceeds by ordinal and an unready Pod can prevent later Pods from updating. If a bad template has been reverted but rollout remains blocked, the StatefulSet documentation describes deleting the Pod created from the bad template so the controller can recreate it from the corrected template. Confirm the Pod, ordinal, template, and storage implications before taking that action; do not improvise a data-changing recovery outside the application’s recovery procedure.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

StatefulSet update controls are version-sensitive. The current documentation marks maxUnavailable beta since Kubernetes v1.35 and a Recreate strategy alpha since v1.37, disabled by default behind a feature gate. Do not assume either is available or enabled on an incident cluster: verify its Kubernetes version and feature gates. The official StatefulSet guide explains ordered rollout and rollback behavior.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Why a StatefulSet rollout can get stuck

In an ordered rollout, the controller waits for the relevant Pod to become ready before advancing. A Pod that cannot schedule, mount its claim, start, pass readiness, or serve the required application behavior can stop progress at its ordinal. Inspect the failing Pod and its Events, then correlate the ordinal with the claim, volume, and application logs. Reverting the template alone may not clear a Pod already created from the bad template; use the documented recovery behavior and the application’s data-recovery procedure before deleting or replacing it.

Cleanup has its own storage and termination caveats: scaling down or deleting a StatefulSet does not delete its associated volumes, and deleting the StatefulSet does not guarantee ordered, graceful Pod termination. Plan volume lifecycle and application shutdown explicitly rather than treating controller deletion as a complete cleanup or recovery action.

What to improve after the incident

  • Make rollout risk explicit. Review update budgets, capacity, readiness conditions, and whether a canary or staged rollout is appropriate for the workload and controller.
  • Observe the layer that failed. Alert on application health and request behavior as well as Pod and controller status; for DaemonSets, monitor coverage across eligible nodes, and for StatefulSets, correlate ordinals with claims and storage events.
  • Validate recovery paths. Test rollback and application recovery procedures, including any data migrations or external configuration changes that a controller rollback cannot undo.
  • Use disruption controls for their intended scope. A PodDisruptionBudget can govern certain voluntary disruptions, but it is not a limit on a Deployment or StatefulSet’s own rolling upgrade. Do not treat it as a complete rollout safety rail; see Kubernetes’ Disruptions documentation.
  • Record version and configuration. Preserve the cluster version, workload templates, strategy settings, selectors, node labels, and relevant feature gates with the incident timeline so later analysis reflects the system that actually ran.

Kubernetes’ workload management guidance provides the broader context for managing rollouts across controllers.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.