Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

If a Kubernetes HorizontalPodAutoscaler (HPA) keeps more replicas than the latest metric seems to require, check its scale-down stabilization window and policies first. Then inspect HPA status and Events, verify the relevant metrics API, confirm the replica floor and workload ownership, and—when scaling on CPU—review resource requests and pod readiness. Scale-to-zero has separate requirements.

Why does an HPA keep more replicas after demand falls?

HPA does not necessarily set the target to the replica count implied by the newest metric sample. It calculates recommendations, records them, and applies behavior rules before updating the scalable target. As the Kubernetes HPA concepts documentation puts it: “Finally, right before HPA scales the target, the scale recommendation is recorded.”

Scale-down stabilization holds recent high recommendations

The documented default scale-down stabilization window is 300 seconds (five minutes). During that window, HPA uses the highest recent recommendation, so a recent high value can keep the desired count above what the newest low metric would suggest. This is intentional protection against a brief dip causing a rapid reduction.

The cluster-wide default can be changed with the controller-manager flag --horizontal-pod-autoscaler-downscale-stabilization. An individual HPA can set spec.behavior.scaleDown.stabilizationWindowSeconds. The autoscaling/v2 API reference permits values from 0 to 3600 seconds. A zero window removes the history-based delay but also removes that protection against metric fluctuation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Scale-down policies can limit or disable reductions

After calculating the desired count, HPA applies scale policies. The documented default scale-down policy allows removal of all replicas above the minimum within its 15-second policy period. A custom policy can make reductions more gradual; when multiple policies exist, selectPolicy determines which applies. For example, Min selects the smallest permitted change, while Disabled turns off scaling in that direction.

Inspect the live HPA’s spec.behavior.scaleDown, not just a file in source control. The same API reference documents the default cluster-wide tolerance of 10% for small metric variations when no tolerance is set; this is a configuration default, not a guarantee that every cluster uses it.

What should I check first in HPA status and Events?

Start by identifying the HPA’s actual target and observed state. A DaemonSet is not a valid HPA target because it does not expose the required scale subresource; supported targets include scalable workloads such as Deployments and StatefulSets.

  1. kubectl get hpa — identify the HPA and compare its current and desired replica counts.
  2. kubectl describe hpa <name> — inspect the target reference, minimum and maximum replicas, configured metrics and targets, Conditions, and recent Events.
  3. kubectl get deployment <name> or kubectl get statefulset <name> — confirm which workload is being scaled and its current replica state.

The Kubernetes HPA walkthrough explains the main Conditions:

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • AbleToScale indicates whether HPA can fetch or update the scale, including whether backoff-related conditions affect scaling.
  • ScalingActive indicates whether HPA is enabled and can calculate a desired scale. A false value commonly points to a metrics problem.
  • ScalingLimited indicates that the desired scale was constrained by a minimum or maximum boundary.

Events help distinguish metric retrieval or conversion failures from scale access, backoff, or min/max constraints. Do not diagnose from the current metric alone: behavior rules may still hold or limit the resulting recommendation.

How can missing metrics block scale-down?

Find out which metrics the HPA uses and whether each corresponding API can serve them. Per-pod CPU and memory resource metrics come from metrics.k8s.io, commonly provided by metrics-server. Custom and external metrics use custom.metrics.k8s.io and external.metrics.k8s.io, typically provided by metrics adapters. The aggregation layer and API registrations must be available for HPA to retrieve them; see the Kubernetes metrics API documentation.

Review the metric and target shown by kubectl describe hpa <name>, then look for Events reporting retrieval or conversion errors. HPA handles missing data conservatively: when considering a possible scale-down, it assumes pods with missing metrics consume 100% of their target. For an HPA configured with multiple metrics, it chooses the largest desired replica count. If one metric cannot be converted to a desired count while another valid metric recommends scaling down, HPA skips that scale-down.

Consequently, low CPU does not guarantee a reduction if a custom or external metric is unavailable. Resolve the API, adapter, or metric-query problem before loosening scale-down safeguards.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Could the minimum replica count or another writer be responsible?

Check the replica floor

HPA will not reduce the target below minReplicas. Compare that value in the live HPA with the workload’s current count. If ScalingLimited indicates a lower-bound constraint, change the minimum only if the workload’s availability requirements permit it.

Look for automation that resets replicas

When HPA manages a Deployment or StatefulSet, repeatedly applying a workload manifest that sets a fixed spec.replicas can reset the target and cause replica thrashing. Kubernetes recommends removing spec.replicas from the workload manifest when HPA manages scaling. Also check deployment automation and other controllers that may write the scale subresource.

Why can CPU-based scaling respond less than expected?

CPU utilization is calculated relative to the pods’ CPU resource requests. If a relevant container lacks a CPU request, utilization for that metric cannot be calculated as expected. Check requests on the actual pods, not only the HPA target.

HPA also treats not-yet-ready pods and startup CPU samples specially, and it accounts conservatively for missing metrics. Those safeguards can dampen the size of a scale change. The documented controller-manager defaults for CPU startup handling are a 30-second initial readiness delay and a five-minute CPU initialization period; confirm the cluster’s actual settings because these are cluster-wide controller options.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A startup probe or readiness probe that reflects when the application has finished its startup CPU spike can help keep that spike from distorting CPU-based scaling. A probe should represent the application’s actual startup or readiness state, rather than being adjusted solely to force a desired replica count.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How should I troubleshoot HPA scale-to-zero?

Treat zero replicas as a separate case. Current Kubernetes documentation describes HPA scale-to-zero with minReplicas: 0 and object or external metrics. CPU and memory resource metrics cannot trigger scaling from zero because no pods remain to provide those metrics.

The Kubernetes v1.37 announcement says HPAScaleToZero is enabled by default in v1.37 and describes the ScaledToZero condition, which helps distinguish an HPA-managed zero-replica state from a manually paused workload. Check that condition, the external or object metric, and feature support in both kube-apiserver and kube-controller-manager. During a version-skewed upgrade, the announcement advises waiting until both components support the feature before using minReplicas: 0.

If the metric adapter cannot return the required metric, HPA may report ScalingActive=False with a reason such as FailedGetExternalMetric. Use the condition and Events to determine whether zero replicas reflect HPA behavior or a metric/API failure.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Which settings should I change?

Make a change only after matching the observed symptom to the relevant control. Kubernetes documentation establishes the available behavior, not a universally correct setting for every workload.

Setting or check What it controls Trade-off or use
stabilizationWindowSeconds How long recent recommendations affect scale-down; API range is 0–3600 seconds. A shorter window responds sooner after a dip but provides less protection against transient drops.
scaleDown.policies and selectPolicy How quickly the replica count may decrease, or whether that direction is disabled. More restrictive policies reduce replicas gradually; permissive policies allow faster reduction.
minReplicas The lowest replica count HPA is allowed to maintain. Lowering it allows fewer replicas but must remain compatible with availability needs.
Metrics API and adapter health Whether HPA can retrieve and convert each configured metric. Fix an unavailable metric source rather than weakening scale-down behavior to mask the failure.

Before changing values, compare the live HPA spec and status with the Kubernetes version and controller-manager settings actually deployed. Current documentation describes defaults; an older release or altered controller flags can behave differently.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.