The settings that matter most are the scaling signal and target, replica limits, scale-up and scale-down behavior, and the resource requests that underpin both pod scaling and node placement. Tune them as one capacity path: a HorizontalPodAutoscaler (HPA) can add replicas, but if the cluster has no room to schedule them, node autoscaling must supply capacity too.
Which settings should you tune first?
- Choose a signal that reflects the workload bottleneck. Use CPU or memory utilization when those resources limit the application and its resource requests are credible. Otherwise, consider an application-relevant metric such as requests per second, work per pod, or queue depth.
- Set a target and test it against service behavior. A metric must reflect work that additional replicas can actually absorb. Validate the target against latency, saturation, queueing, and replica startup data rather than assuming a technically available metric is useful.
- Set minimum and maximum replicas deliberately. The minimum represents warm capacity and availability; the maximum is a capacity and spend guardrail. Choose both based on startup time, expected demand, per-replica capacity, disruption tolerance, and the service’s SLO.
- Shape scale-up and scale-down behavior. Policies limit the rate of replica changes, while stabilization windows smooth recommendations. Check how those choices perform against bursts, transient dips, and scaling delay.
- Include node autoscaling in the plan. If existing nodes cannot fit new pods, the node autoscaler must provide schedulable capacity. Review pod requests and node-group policies alongside HPA settings.
These are Kubernetes concepts, not universal tuned values. Defaults and supported settings can vary by Kubernetes release, cluster configuration, managed service, and node-autoscaler implementation.
Choose the right metric and target
HPA can use resource metrics, per-pod metrics, object metrics, and external metrics. Kubernetes examples include transactions per second, ingress hits per second, queue length, and load-balancer QPS. The best choice depends on which measurement tracks the workload the replicas can handle.
- CPU or memory utilization: Useful when the resource is a genuine capacity constraint and resource requests provide a meaningful baseline. Utilization-based targets are calculated relative to requested resources.
- Requests per second or work per pod: Can be a closer indicator for request-driven services when replica capacity corresponds to that workload.
- Queue depth: Can signal accumulating work, but confirm that more replicas can process the queue and that the metric is timely.
- Object or external metrics: May represent load outside an individual pod. Check freshness, availability, and how the signal maps to replica demand.
With multiple metrics, HPA selects the largest desired replica count among the metrics it can calculate. If a metric fails while available metrics indicate scaling down, HPA skips that scale-down. This behavior makes metric quality and availability important to both responsiveness and stability. See the Kubernetes HPA documentation and the HPA v2 API reference.
#1 Best Overall
Make resource requests credible
Requests connect HPA decisions to scheduling. For utilization-based scaling, utilization is measured relative to requested CPU or memory. Requests also help the scheduler decide where pods fit and affect node autoscaling and consolidation. If requests do not approximate workload needs, both controllers work from a misleading capacity picture.
- Requests that are too low can make utilization look higher relative to the requested amount and can undermine capacity planning. Kubernetes also cautions that requests that are too low can make new-node provisioning ineffective.
- Requests that are too high can make available node capacity appear unavailable to new pods and can block node consolidation.
- Use measured workload behavior to set requests, then review their effects on HPA recommendations, pending pods, and node utilization.
For the relationship between pod scaling and node capacity, see Kubernetes Node Autoscaling.
Set replica bounds around availability and spend
minReplicas and maxReplicas define the HPA’s operating range. There is no evidence-based universal count: the right bounds depend on the application, its SLO, startup time, replica capacity, expected traffic, and tolerance for disruption.
- Minimum: A higher minimum keeps more warm capacity available, which can help absorb demand before new pods start. It also retains more replicas when demand is low.
- Maximum: A higher ceiling allows more scale-out when capacity is needed, but can permit greater resource use. A low ceiling constrains spend but may leave demand unmet if the service needs more replicas.
Check whether the maximum is achievable on the available node groups. An HPA can request more pods than the cluster can currently schedule; node provisioning and its constraints are part of the outcome.
Recommended Free Tools
Balance scale-up speed, scale-down retention, and tolerance
Scale-up policies constrain how quickly replicas can increase. More aggressive behavior may bring capacity online sooner, but can create excess replicas or churn; conservative behavior may delay capacity during a surge. Actual response also depends on metric evaluation, pod readiness, and—when nodes must be added—node provisioning.
Scale-down policies and stabilization windows govern how quickly replicas can be removed. Retaining capacity through short demand dips can preserve headroom if load rebounds, while faster reduction can lower ongoing pod use. The tradeoff is specific to the workload’s traffic pattern and scaling delay.
Rank #3
The current Kubernetes HPA API reference documents these defaults when behavior is unspecified: a 0-second scale-up stabilization window and a 300-second scale-down stabilization window. It documents a 10% cluster-wide default tolerance when no tolerance is set. These are documented defaults, not recommendations for every workload; verify the Kubernetes release and cluster configuration in use. A tolerance filters small changes around the target: lower tolerance can react to smaller deviations, while higher tolerance can reduce churn at the cost of waiting longer to adjust.
Understand why pods can remain pending after HPA scales up
HPA adjusts workload replicas; node autoscaling provisions or consolidates cluster nodes. They are complementary controllers, not interchangeable settings. If existing nodes cannot fit the additional pods, a node autoscaler must find eligible capacity and provision nodes before those pods can schedule.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Repair Windows errors before they cause bigger problems3Fix the driver behind crashes, sound loss and screen glitchesThe Cluster Autoscaler project’s FAQ describes separate contributors to this delay: HPA reaction, node-autoscaler reaction, and node provisioning. It lists up to 10 seconds before scale-up is considered and 10 minutes before a node is considered unneeded as documented Cluster Autoscaler defaults. The FAQ also reports 3 to 4 minutes on GCE from a Cluster Autoscaler request until pods can be scheduled on new nodes, and about 5 minutes for the total HPA-plus-Cluster-Autoscaler flow under its described assumptions. The GCE and total-flow figures are project-specific operational observations, not guarantees for other providers or clusters. Check the deployed version, flags, provider, and environment in the Cluster Autoscaler FAQ.
When investigating pending pods, check whether requested CPU and memory fit available nodes, whether pod constraints permit placement on the node groups that can scale, and whether node provisioning is still in progress. A pod count increase alone does not show that usable capacity has arrived.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Choose node-group behavior for schedulability and cost
When multiple node groups are eligible, the node autoscaler may need to choose among them. The Cluster Autoscaler FAQ describes strategies including most-pods, least-waste, least-nodes, price, and priority; availability depends on the implementation and provider.
Evaluate a strategy against whether constrained pods can schedule, how much CPU and memory remain unused after scale-up, the number of nodes required, and the price of eligible nodes. A cost-oriented choice is only useful if the selected capacity can actually run the workload.
Best Value
Validate the whole control loop
Do not judge autoscaling from replica count alone. Observe how the signal changes, when HPA recommendations change, when new pods become ready, whether pods remain pending, when nodes arrive, and how latency and spend respond. Kubernetes HPA periodically evaluates metrics and accounts for not-yet-ready pods and missing metrics; initialization and readiness handling can affect CPU-based recommendations.
- Compare the chosen metric with latency, saturation, and queueing during both normal traffic and bursts.
- Track time from demand change to HPA recommendation, pod readiness, and—if required—schedulable node capacity.
- Look for overprovisioning and churn after demand falls, as well as insufficient headroom when demand rebounds.
- Revisit requests if HPA utilization, pod placement, or node consolidation does not match observed workload behavior.
- Confirm API support, defaults, and node-autoscaler behavior against the deployed Kubernetes version and provider.
The Kubernetes node-autoscaling documentation describes the intended balance: “If configured correctly, this pattern ensures that your application always has the Node capacity to handle load spikes if needed, but you don’t have to pay for the capacity when it’s not needed.” That outcome depends on configuration and workload behavior; it is not a guarantee that any particular settings will meet an SLO or cost target.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

