Recommended Free Tools
Repeated scale-up and scale-down cycles—often called “thrashing” or “flapping” in Kubernetes documentation—usually mean the HorizontalPodAutoscaler (HPA) is reacting to a noisy or misleading signal, or that another system is also changing the workload’s replica count. Start by tracing who controls replicas, then follow each HPA recommendation back to its metric, target, and data pipeline before tuning scaling behavior.
This guide focuses on Kubernetes HPA. Its documented API is autoscaling/v2, but defaults and available fields depend on the cluster’s Kubernetes version, controller-manager configuration, and managed-service behavior. Verify those details before changing settings.
1. Confirm who is changing the replica count
First distinguish an HPA decision from an external change. A GitOps controller, operator, deployment tool, or person may also write the workload’s replica count, causing the observed value to differ from what the HPA wants.
- Run
kubectl get hpato identify the HPA and see its target, current and desired replica counts, and reported metrics. - Run
kubectl describe hpa <hpa-name>to inspect metric readings, conditions, and recent events. - Check the target workload’s scale state and audit the systems that manage its configuration. Confirm whether any of them write replicas independently of the HPA.
The HPA inspection commands show its view of scaling; detecting other writers requires checking the tools and controllers installed in your cluster.
#1 Best Overall
2. Trace the recommendation to its metric and target
For every HPA metric, identify the target value and compare it with the metric value the HPA reports. The HPA calculates a desired replica recommendation from the metric-to-target ratio. With multiple metrics, it calculates a recommendation for each and chooses the largest desired count. One metric can therefore request more replicas even when another appears to justify fewer.
Check that each metric’s units, labels, selector, aggregation, and observation time match your intent. If a metric is missing or cannot be fetched, the HPA may be unable to apply a scale-down recommendation when another available metric recommends scaling down. Do not adjust targets until you know which metric is driving the recommendation and whether its data is valid. See the Kubernetes HPA documentation and the autoscaling/v2 API reference.
3. Verify the metrics pipeline
Check the API serving the exact metric requested by the HPA, not merely whether some metrics are visible elsewhere.
- CPU and memory: Verify readings through the resource Metrics API and that metrics-server is collecting and aggregating kubelet data. Kubernetes describes this as the basic pipeline for Pod and node resource metrics.
- Custom or external metrics: Confirm the corresponding metrics API and adapter are available, the mapping is correct, and the requested metric has the expected scope and labels.
- Intermittent gaps: Compare metric timestamps and adapter/API responses with HPA conditions and events. A working resource Metrics API does not establish that custom or external metrics are available.
See Kubernetes resource metrics pipeline documentation for the resource-metrics path.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
4. Check startup CPU and readiness effects
A new Pod’s startup CPU burst or readiness changes can distort the signal the HPA uses, especially when CPU utilization is calculated against resource requests. Check whether requests represent the workload’s expected consumption and whether probes mark the Pod ready only after its startup behavior has stabilized.
Kubernetes documents two controller-manager settings relevant to CPU samples: --horizontal-pod-autoscaler-cpu-initialization-period, with a documented default of 5 minutes, and --horizontal-pod-autoscaler-initial-readiness-delay, with a documented default of 30 seconds. These are cluster-wide controller settings, not per-HPA settings; confirm the actual values and behavior for your cluster before changing them.
Rank #3
5. Compare signal timing with workload timing
Write down the metric scrape and aggregation interval, HPA reconciliation cadence, time needed for new Pods to become useful, and typical duration of demand bursts. A short-lived threshold crossing can prompt one recommendation, followed by a reversal when the next observation arrives. A scale-down that begins before replacement capacity is ready can also make the service appear unstable.
Use service indicators alongside replica counts while diagnosing: queue depth, latency, saturation, and error rate help show whether a replica change followed real demand or only a transient reading. Change one relevant control at a time so you can see whether the service outcome improves.
Free tools Windows power users keep installed
One-click scans. No signup required.
6. Choose scale-down smoothing without blunting scale-up
Stabilization windows
spec.behavior.scaleDown.stabilizationWindowSeconds makes the HPA consider recent recommendations before applying a decrease. The documented default is 300 seconds (five minutes); within that window the HPA uses the highest recommendation, buffering a temporary load dip. The documented scale-up stabilization default is 0 seconds, so scale-up can proceed immediately subject to applicable policies.
Rank #4
A stabilization window smooths decisions; it does not repair a bad metric or resolve competing replica writers. If short dips are causing churn, a longer downscale window may help, but it also keeps capacity running longer after demand falls.
Directional scaling policies
HPA behavior supports separate policies for scale-up and scale-down, including limits on how much the replica count may change over a period. Policies can also disable scaling in one direction. The API reference describes a default scale-up policy allowing at most doubling replicas or adding four Pods over a 15-second period; actual behavior depends on release and configuration.
Restrict scale-down when rapid decreases cause churn. Limit scale-up only when the workload can tolerate a slower response: too much smoothing in that direction can leave capacity short and increase latency.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Scan for outdated or missing drivers - takes under a minute3Clear out junk files and repair common Windows errorsTolerance around the target
The documented default metric tolerance is 10%, unless cluster-wide or per-HPA configuration overrides it. Tolerance creates a band around the target in which small variations do not change the desired count. Per-direction tolerance support and API fields depend on Kubernetes version and feature availability, so check the cluster’s release before using them.
For the behavior and version-sensitive details, consult the HPA API reference and HPA documentation.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.7. Use the symptom to pick the next check
| What you see | What to inspect or test |
|---|---|
| Metric readings hover around the target | Verify units and aggregation first; then assess whether tolerance or a longer scale-down stabilization window would reduce reversals without harming capacity. |
| Scale-up follows startup CPU, then scale-down follows readiness | Inspect startup and readiness probes, resource requests, and CPU initialization handling. |
| HPA status and workload replicas disagree | Compare HPA status and conditions with the target’s scale state; check for another system writing the replica field. |
| Several metrics appear to compete | Inspect every metric separately. The largest successful desired-replica recommendation wins. |
| Metric data intermittently disappears | Check availability and mappings for the specific resource, custom, or external metrics API the HPA uses. |
| Unexpected scaling to zero | Check the exact Kubernetes release and metric type. In Kubernetes v1.37, scaling to zero is Beta and enabled by default for object and external metrics, not CPU or memory alone. Consider whether a durable queue or other buffering layer can preserve work during cold starts. |
The v1.37 scale-to-zero behavior is release-specific; verify it against your cluster’s version and configuration. See Kubernetes HPA concepts.
8. Compare a proposed configuration against service outcomes
There is no universal best HPA setting. When comparing configurations, evaluate how quickly each responds to a real load increase, how much transient noise it tolerates, how quickly it removes capacity after load falls, and whether its metric remains available when replicas are starting or at zero. Then compare the effect on latency, queue delay, errors, and infrastructure cost.
Keep scale-up responsiveness and scale-down smoothing as separate decisions. A longer downscale window can dampen short dips while preserving a faster upward response; verify the result against actual workload behavior rather than treating a stable replica graph as proof of healthy service.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

