Use a Kubernetes HorizontalPodAutoscaler (HPA) to scale a workload’s replica count, and configure its minimum and maximum replicas, scale-rate policies, and stabilization windows around your service’s tested capacity. HPA does not have a setting literally named “cooldown”: stabilization smooths recommendations, while rate policies limit how quickly the replica count can change.
What an HPA controls—and what it does not
An HPA periodically reads workload metrics and adjusts the desired replica count through the workload’s scale subresource. It is a control loop, not an instantaneous reaction; Kubernetes documents a default controller sync period of 15 seconds. A change in replicas may take longer to affect serving capacity because new Pods still need to be scheduled, started, and become Ready. See the HPA concepts documentation for the control-loop details.
HPA scaling and node autoscaling are separate layers. HPA requests more or fewer workload Pods. Node autoscaling changes cluster infrastructure when capacity needs change. If the cluster cannot schedule the replicas HPA requests, those Pods may remain pending until capacity is available; increasing the HPA maximum alone does not add nodes.
Set bounds from capacity, not from a generic example
Use the stable autoscaling/v2 API and target the Deployment or other scalable workload that should change. The two essential bounds are minReplicas and maxReplicas; the maximum must not be lower than the minimum. The HorizontalPodAutoscaler API reference defines these fields, but Kubernetes does not provide universally safe replica counts.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →#1 Best Overall
Choose the bounds using evidence about your service and its dependencies. The minimum is the floor you are willing to keep running under low demand. The maximum is the most replicas the service and its environment can safely support, considering tested throughput, latency objectives, resource budgets, and downstream limits such as database connections or API quotas. Validate the maximum under load: replicas that exceed a dependency’s capacity can amplify overload rather than relieve it.
- Establish how much traffic or work one healthy replica can handle while meeting the service’s latency objectives.
- Check that the maximum replica count fits the available resource budget and does not overwhelm dependencies.
- Choose a minimum consistent with the service’s baseline demand and acceptable response time when load rises.
- Revisit both bounds when workload behavior, resource sizing, or dependency capacity changes.
Choose metrics that reflect the bottleneck
HPA needs metrics that change predictably when replicas are added or removed. For CPU or memory resource metrics, the cluster needs the resource metrics API, commonly supplied by Metrics Server. CPU utilization is calculated relative to CPU requests, so define appropriate CPU requests if you use a CPU-utilization target; without them, that percentage cannot provide a meaningful capacity signal.
Rank #2
HPA can use multiple metrics. When they imply different replica counts, it uses the largest desired count. This is useful when any one of several bottlenecks should trigger scale-up, but it also means a high metric can keep the workload above the count suggested by the others. Metric availability affects decisions: missing metrics and Pods that are not yet Ready are handled conservatively, and a metric error can prevent a scale-down recommendation. Check the HPA’s status and conditions when its actions do not match expectations. The Kubernetes HPA documentation describes these calculations and readiness considerations.
Startup behavior matters particularly for CPU-based scaling. Kubernetes documents default values of 30 seconds for the initial readiness delay and five minutes for the CPU initialization period. These windows affect how startup metrics are treated; make sure Pod readiness represents actual ability to serve, and account for initialization behavior when diagnosing unexpected scaling.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Use rate policies to limit replica changes
In autoscaling/v2, behavior.scaleUp and behavior.scaleDown let you control each direction independently. A Pods policy limits the number of replicas that may change during a period; a Percent policy limits the change as a proportion of the current replica count. Each policy specifies periodSeconds, the interval over which its allowed change is evaluated.
If a direction has multiple policies, the default selection allows the largest change permitted by any policy. Set selectPolicy: Min when you want the strictest of the configured limits to apply. Set selectPolicy: Disabled to turn scaling off in that direction. These controls are documented in Kubernetes’ guide to configurable scaling behavior.
Rank #4
Choose policy types and periods according to the risk of changing capacity. A faster or larger scale-up can help absorb rising demand but may put pressure on downstream services. A stricter limit protects those services but can leave the workload short of capacity for longer. For scale-down, rate limits can prevent a large abrupt reduction even when recommendations fall quickly. Set policies based on load tests and service behavior rather than copying an unvalidated replica count or interval.
Use stabilization windows to smooth noisy recommendations
A stabilization window and a rate policy solve different problems. The window smooths the HPA’s recommendations; the policy caps how much the replica count can change over a specified period. Kubernetes documents a default scale-down stabilization window of 300 seconds (five minutes). During that window, the controller uses the highest recent desired replica recommendation to avoid scaling down in response to a brief dip. Kubernetes documents no default scale-up stabilization window.
For scale-down, the documented five-minute default is a reasonable starting point unless workload response or cost requirements justify changing it. A longer window retains capacity through short-lived drops but keeps more resources running; a shorter one can release resources sooner but is more exposed to transient low readings. Validate the chosen window against startup time, queueing behavior, and service-level objectives. The documentation defines how the mechanism works, not a universally correct duration.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Build and validate the HPA configuration
- Identify the target workload. Choose the scalable resource whose replicas should change, and confirm that it can serve additional traffic when more Pods become Ready.
- Select a usable metric. Confirm that the metric is available to HPA and that it responds in a predictable way to adding or removing replicas. For CPU utilization, set appropriate CPU requests.
- Set the replica bounds. Configure
minReplicasandmaxReplicasfrom tested service capacity, dependency limits, latency objectives, and budget. Ensure the maximum is not below the minimum. - Configure each direction separately. Under
behavior.scaleUpandbehavior.scaleDown, choose Pods or Percent policies and theirperiodSeconds. SetselectPolicy: Minwhen the strictest of multiple limits must apply. - Choose stabilization deliberately. Use the scale-down window to filter short-lived drops; consider scale-up stabilization only if delaying response to an increase is acceptable. Do not treat a rate policy as a substitute for a window.
- Check metrics and readiness. Verify the resource metrics API or the relevant alternative metrics source, and ensure readiness accurately indicates serving ability.
- Test both directions and inspect status. Exercise rising demand, sustained load, brief spikes, and falling demand. Observe HPA status and conditions, Pod readiness, scheduling, and downstream health; adjust bounds and behavior if the observed response violates your capacity or latency goals.
Plan cluster capacity separately
HPA can request replicas beyond the capacity of the nodes currently available. If Pods cannot be scheduled, investigate cluster capacity and node autoscaling rather than assuming the HPA has failed. Node autoscalers use Pod resource requests to make placement and capacity decisions, so requests should reflect realistic workload needs. HPA bounds protect workload-level scaling; they do not guarantee that nodes or dependencies can support every replica within those bounds.
When scale-to-zero is appropriate
Zero replicas require a metric that can trigger activation while no workload Pods are running. CPU and memory resource metrics cannot do that because they depend on running Pods. In Kubernetes v1.37, scale-to-zero is documented as beta and is available for object or external metrics, not resource metrics. The v1.37 documentation says minReplicas: 0 requires at least one object or external metric and the HPAScaleToZero feature gate enabled in both kube-apiserver and kube-controller-manager; see the Kubernetes v1.37 announcement.
Confirm the cluster version and feature-gate state before relying on this behavior; the v1.37 status should not be assumed for older clusters. Scale-to-zero is appropriate only when an external signal can wake the workload and its activation delay is acceptable.
Free tools Windows power users keep installed
One-click scans. No signup required.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

