Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Start by comparing the HPA’s desired replica count with the workload’s actual count. If the desired count is still high, check the stabilization window, metric calculations, and replica limits. If the desired count is lower than the workload’s count, investigate scale access, events, and other controllers or manifests that may be writing replicas.

1. Find out whether the HPA wants fewer replicas

Inspect the HPA and its target before changing settings. Run kubectl get hpa to see the HPA’s current and desired replicas and reported metrics, then use kubectl describe hpa <name> for detailed status, conditions, and events. Inspect the target workload’s actual replica count as well.

Compare the HPA’s desired count with the target’s actual count:

  • Desired count is still high: The HPA’s calculation has not produced a lower target. Continue with its timing, bounds, and metrics.
  • Desired count is lower than the target’s actual count: The HPA has calculated a reduction but it may be unable to apply it, or another writer may be resetting the count.

Check the HPA conditions. AbleToScale reports whether it can fetch or update scale, including whether backoff is preventing an update. ScalingActive indicates whether scaling is active. ScalingLimited indicates that the desired scale was constrained by configured bounds.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall

2. Check when the HPA can act on a lower recommendation

HPA is a periodic control loop, not an immediate reaction to a metric change. Kubernetes documents a default controller sync period of 15 seconds, but reconciliation can take longer depending on metrics and control-plane conditions.

Scale-down has a separate default stabilization window of 300 seconds (five minutes). During that window, the controller uses the highest recent replica recommendation. A short-lived high recommendation can therefore keep the desired count above the newest recommendation until it falls outside the window.

Inspect spec.behavior.scaleDown.stabilizationWindowSeconds, the scale-down policies, and selectPolicy. A policy can limit how quickly replicas are removed; selectPolicy: Disabled disables scaling in that direction. If considering a shorter window or a less restrictive policy, weigh the desired response time against pod startup time, latency sensitivity, and the risk of replica churn under brief metric dips.

3. Verify the target and replica limits

Check the HPA’s scaleTargetRef, minReplicas, and maxReplicas. The target must implement the scale subresource, and its reference and labels determine which pods are selected for metrics.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • If the workload is already at minReplicas, the HPA cannot reduce it further.
  • If ScalingLimited is true, read its reason and compare the desired count with the configured bounds.

4. Validate every metric configured on the HPA

Check that each metric source is available and that its name, selector, and target configuration match the HPA. The relevant API depends on the metric type: resource metrics use metrics.k8s.io, custom metrics use custom.metrics.k8s.io, and external metrics use external.metrics.k8s.io. Metrics Server is a common provider for resource metrics; custom and external metrics depend on the cluster’s adapter and metrics pipeline.

Do not dismiss a metric error just because another metric looks healthy. With multiple metrics, HPA normally uses the largest replica recommendation. If one metric cannot be converted to a desired replica count while another available metric suggests scaling down, Kubernetes skips the scale-down. Resolve the failing metric path or confirm that the metric should remain configured before attributing the behavior to stabilization.

5. Check pod samples, CPU requests, and readiness

For a CPU utilization target, utilization is evaluated relative to CPU requests. Verify that relevant containers have CPU requests; sidecars matter too unless the HPA uses a container resource metric. Without suitable requests, CPU utilization is not a sound basis for the expected calculation.

Incomplete or unsettled pod data can also make a reduction more conservative. For scale-down recalculation, the controller treats pods with missing metrics as consuming 100% of the target metric. CPU samples from initializing or not-yet-ready pods may be set aside under the controller’s readiness rules. Check selected pods’ readiness and restarts, along with metric freshness and whether the metrics pipeline has samples for every selected pod.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

6. Look for another writer changing the replica count

If the HPA’s desired count is lower than the workload’s actual count, or replicas keep returning after a change, inspect workload manifests and automation for a hard-coded spec.replicas. A rollout, GitOps reconciliation, operator, or human action may write a replica count after the HPA updates it.

Kubernetes recommends omitting spec.replicas from Deployment or StatefulSet manifests managed by an HPA. Applying a manifest that sets the field can reset the live count and cause the observed flapping.

Can this HPA scale to zero?

Kubernetes v1.37 documentation describes scale-to-zero as a beta feature enabled by default in that version. It applies to object or external metrics with minReplicas: 0, not CPU or memory resource metrics, which require running pods. The feature gate must be enabled on both the API server and controller manager. Check the cluster’s actual Kubernetes version and feature-gate configuration; scale-to-zero does not remove the ordinary minimum-replica limit for a resource-metric HPA.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.