Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

A Kubernetes rolling update limits how many Deployment Pods can be unavailable or added at once; it does not guarantee that replacement Pods can serve real requests or that every component on the request path has stopped sending traffic to terminating Pods. To find the cause of 503s, correlate the error window with Pod health, Service endpoints, rollout settings, shutdown behavior, and the proxies or load balancers handling traffic.

What a rolling update controls—and what it does not

A Deployment using the RollingUpdate strategy replaces Pods gradually. Its maxUnavailable setting limits how many replicas may be unavailable during the update, while maxSurge limits how many extra Pods may be created. Kubernetes documents defaults of 25% for both; percentage values round down for maxUnavailable and up for maxSurge. These are controller limits, not proof that an application is serving correctly or that every traffic component has observed an endpoint change. See the Kubernetes rolling update guide and Deployment documentation.

Replica counts can conceal a capacity problem, especially in a small Deployment: rounding a percentage can produce a different practical allowance than the percentage suggests. During the error window, record desired, updated, ready, and available replica counts, and inspect the actual rollout settings rather than assuming the defaults apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Check whether new Pods are really ready to serve

A Pod can be running without being ready. Readiness determines whether it is eligible for Service traffic; a failed readiness probe leaves the container running but marks the Pod not ready. A liveness failure, by contrast, can restart the container. If readiness becomes true before the application can handle routed requests—for example, while it is still warming caches or waiting on a dependency—the rollout can direct traffic to a process that is not operationally ready. If liveness is used for a condition that should only keep traffic away, it may instead cause unnecessary restarts.

Review the probe path and its meaning against the requests that are failing. A successful probe should indicate that the application can handle the traffic it will receive, not merely that the process or a lightweight endpoint responds. Startup probes can postpone readiness and liveness checks until initialization succeeds, which is useful when startup takes longer than normal operation. Check probe timing, thresholds, and the application’s actual warm-up behavior together. Kubernetes explains these distinctions in Configure Liveness, Readiness and Startup Probes.

Correlate 503 timestamps with Service endpoints

During the rollout, inspect the Service’s EndpointSlices and compare endpoint conditions with the times 503s occurred. In particular, note ready, serving, and terminating. These states help show whether endpoints were becoming eligible, continuing to serve, or shutting down as requests failed. An endpoint can remain represented while a Pod terminates; its conditions and the behavior of the component consuming the endpoint information matter. Consult Kubernetes’ EndpointSlices documentation and its Pod and endpoint termination flow.

Do not infer from a healthy Deployment status alone that every intermediary has converged. Trace the actual request route: Service proxy, ingress or gateway, service mesh, external or cloud load balancer, and client. Kubernetes defines controller and endpoint behavior, but those documents do not establish propagation or connection-draining behavior for a particular provider or implementation. Use that component’s configuration, logs, and telemetry to explain its part of the incident.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Inspect how old Pods shut down

When an old Pod is removed, graceful termination must align with application behavior and traffic handling. Check whether the application stops accepting new work, how it completes in-flight requests, and whether a preStop hook is configured. Compare the termination grace period with the intended cleanup and drain time, and correlate shutdown events with request duration and 503 timestamps.

A graceful shutdown configuration cannot by itself guarantee that external clients or load balancers stop routing at the right moment. The Kubernetes termination-flow documentation describes Pod and endpoint behavior; the timing and effectiveness of draining through ingress, mesh, or provider-managed infrastructure must be verified in the relevant implementation.

Use rollout settings and events to narrow the cause

  1. Confirm the workload and strategy. Verify that the affected workload is a Deployment using RollingUpdate. Record desired, updated, ready, and available replicas along with maxUnavailable, maxSurge, and minReadySeconds.
  2. Inspect Pods during the exact error window. Review readiness status, probe configuration, and Pod events. Identify whether new Pods fail readiness, initialize slowly, or become ready before their handlers and dependencies can serve routed traffic.
  3. Inspect the Service’s EndpointSlices. Correlate endpoint ready, serving, and terminating transitions with request failures.
  4. Follow the request path and shutdown timeline. Check logs and telemetry at the application, proxy, ingress or gateway, mesh, and load balancer layers. Compare shutdown, drain, and request timing rather than attributing the error to the Deployment controller without evidence.
  5. Check rollout conditions and events. A Deployment reports progress through conditions and events. Kubernetes documents a default progressDeadlineSeconds of 600 seconds and a default minReadySeconds of 0; verify the manifest and cluster version before relying on defaults. A failed Progressing condition after the deadline signals a stalled rollout, but does not identify the application-level cause.
  6. Consider a PodDisruptionBudget only if relevant. A PDB governs permitted voluntary evictions through the eviction API. It does not replace Deployment rollout limits or repair failed readiness and incorrect health signaling. See the Kubernetes PodDisruptionBudget API reference.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Why the Deployment can look healthy while users see errors

The Deployment controller evaluates rollout progress and replica availability; a user request depends on more than those signals. A new Pod may pass an inadequate readiness check, a terminating Pod may still be involved in traffic handling, or an intermediary may have different endpoint or connection-draining behavior. The useful diagnosis is the event sequence across the request path—not simply whether the strategy is named RollingUpdate.

Kubernetes defaults and fields can vary by release, so check the API behavior for the cluster’s version. The probe documentation marks probe-level terminationGracePeriodSeconds stable since Kubernetes v1.28; confirm version-specific behavior before applying it. The Kubernetes references explain platform behavior, not provider-specific propagation delays or the cause of any particular cluster’s 503s.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.