Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To scale an application in Kubernetes, change its number of Pods or change the resources assigned to each Pod. Use kubectl scale for an immediate replica-count change, a HorizontalPodAutoscaler (HPA) to adjust replicas from metrics, and a VerticalPodAutoscaler (VPA) to adjust per-Pod resource sizing. If new Pods cannot fit on the current nodes, node autoscaling can add cluster capacity. These mechanisms solve different problems and can be combined.

Choose what needs to change

Kubernetes has two workload-scaling approaches: horizontal scaling changes the number of replicas, while vertical scaling changes the CPU and memory assigned to existing replicas. Node autoscaling changes the cluster’s node capacity so that Pods that cannot be scheduled may have somewhere to run. See the Kubernetes overview of workload autoscaling.

Approach What changes What triggers the change What to consider
Manual horizontal scaling Replica count Your kubectl scale command Useful for a deliberate, immediate change; a later controller or configuration change may alter the count.
HorizontalPodAutoscaler (HPA) Replica count Configured resource, custom, or external metrics Requires the appropriate metrics APIs and a scalable workload target.
VerticalPodAutoscaler (VPA) CPU and memory sizing for Pods Observed utilization and other VPA inputs Separate add-on; it adjusts resource sizing rather than adding replicas.
Node autoscaling Node capacity Pods that cannot schedule, plus consolidation opportunities Works from Pod requests and scheduling constraints, not direct post-start usage.
Event- or schedule-driven scaling Usually replica count Events, queue-related signals, or a schedule KEDA supports event-driven scaling and includes a Cron scaler for scheduled scaling.

For an interchangeable, stateless application, a Deployment is a common HPA target. HPA can also target other resources that provide the scale subresource, including StatefulSets. The right target, bounds, metric, and resource settings depend on the application and cluster; Kubernetes documentation does not prescribe universal values.

Change the replica count manually

Use this approach when you want to set a workload’s desired replica count yourself, rather than have a metric-driven controller choose it. Kubernetes documents this example for a Deployment named my-nginx:

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
kubectl scale deployment/my-nginx --replicas=1

Replace my-nginx with the actual Deployment name and 1 with the desired replica count. The command changes the desired number of replicas; it does not itself add nodes if the cluster lacks room to schedule the resulting Pods. See Managing Workloads.

Use an HPA to scale replicas from metrics

An HPA periodically evaluates configured metrics and changes a target workload’s desired replica count to move toward its configured target, while observing its minimum and maximum replica bounds. It is a control loop, not an instantaneous reaction. Kubernetes describes HPA in its Horizontal Pod Autoscaling documentation.

Check the metrics path first

For resource metrics such as CPU and memory, the cluster needs the corresponding metrics API. The resource metrics API is commonly provided by Metrics Server, which is a separate add-on rather than an assumption that every cluster already has the required metrics available. Custom or external metrics require their corresponding metrics APIs and providers. Kubernetes lists monitoring options in Tools for Monitoring Resources.

Set CPU requests when using CPU utilization

CPU utilization targets are calculated relative to resource requests. If a relevant container request is missing, utilization for that Pod is undefined for the metric, so HPA will not act on that Pod for that metric. Requests also influence Pod scheduling, making them important beyond the HPA calculation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose bounds and interpret multiple metrics

Configure a minimum and maximum replica count alongside the metric target. If you configure multiple metrics, HPA evaluates each recommendation and uses the largest recommended replica count, subject to the configured maximum. When metrics are unavailable, a downscale can be skipped while an upscale recommendation may still be followed. Select targets and bounds from workload behavior, service objectives, and capacity constraints rather than treating a generic value as universally safe.

Account for controller timing

Kubernetes documentation gives a default HPA controller synchronization period of 15 seconds and a default downscale stabilization window of 5 minutes. These are controller defaults, not guaranteed end-to-end response times: metric collection, scheduling, Pod startup, readiness, and controller settings also affect how quickly capacity changes become useful.

Avoid competing replica settings

When an HPA owns a Deployment’s replica count, repeatedly applying a manifest that resets spec.replicas to a fixed value can make the desired count thrash or flap. Keep the HPA in control of replica changes, and check the HPA’s status and metric availability if scaling does not behave as expected.

Use a VPA to adjust per-Pod resources

A VPA addresses a different question from HPA: how much CPU and memory should each Pod request or receive? It can adjust resource requests and limits using historical utilization, cluster availability, and real-time signals. VPA is a separately installed add-on represented by a custom resource definition (CRD), and it requires a metrics source. It does not replace HPA’s job of adding or removing replicas. Details are in Kubernetes’ Vertical Pod Autoscaling documentation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Consider VPA when resource sizing is the problem—for example, when existing Pods appear under- or over-provisioned. Consider HPA when the application needs more or fewer replicas as demand changes. If both are used, account for how their decisions interact through resource requests and the workload’s capacity; do not assume they are interchangeable controls.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Scale on events or a schedule

CPU or memory utilization is not the right signal for every workload. For event-driven scaling, including workloads that need to react to an event source or queue-related demand, Kubernetes documentation identifies KEDA. Its Cron scaler supports scheduled scaling. These approaches require their own event or schedule configuration; use the signal that represents the work the application needs to process, rather than assuming CPU utilization will capture every demand pattern. See the Kubernetes workload autoscaling overview.

Add nodes when Pods cannot fit

Workload scaling and node scaling form a chain: HPA can request more Pods as observed demand rises, and node autoscaling can provide capacity for Pods that cannot schedule on existing nodes. Node autoscalers also consolidate underused nodes. They make these decisions using Pod resource requests and scheduling constraints, such as affinity or storage requirements—not the actual resource usage of a Pod after it has started.

That makes accurate requests important to both workload and cluster scaling. If requests are too low, a provisioned node may still be inadequate for the workload’s actual needs. If they are too high, the requests can inhibit consolidation. Kubernetes describes two current node autoscaler implementations sponsored by SIG Autoscaling:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Cluster Autoscaler uses preconfigured node groups.
  • Karpenter uses operator-defined NodePools and has broader node-lifecycle features.

Provider integrations and feature sets differ, so choose based on the specific cluster’s provider support, node configuration model, and lifecycle requirements. See Node Autoscaling.

Troubleshoot scaling that does not take effect

  • The HPA does not change replicas: check that its target supports the scale subresource, that the HPA has the intended minimum, maximum, and metric target, and that the configured metrics API is available.
  • CPU-based scaling seems wrong or incomplete: verify that the relevant containers have CPU requests. Without a request, utilization for that Pod is undefined for the metric.
  • Replicas increase but remain Pending: inspect scheduling constraints and resource requests. If current nodes cannot accommodate the Pods, workload scaling alone cannot create node capacity; node autoscaling may be needed.
  • Replica counts keep changing unexpectedly: check whether an applied manifest is resetting spec.replicas while an HPA also controls the Deployment.
  • Downscaling appears delayed: the default stabilization window is five minutes, and end-to-end timing also depends on metric collection and workload behavior.

Use the HPA’s status and metrics information to distinguish a missing signal from a workload or cluster-capacity constraint. Scaling settings should be validated against the application’s traffic pattern, architecture, Kubernetes version, provider, latency objective, and budget.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.