Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose Kubernetes HPA when CPU, memory, or a metric already available through Kubernetes expresses your workload’s demand. Choose KEDA when you need an event-source scaler or event-driven activation from zero. They are not mutually exclusive: KEDA commonly manages activation between zero and one replica, while HPA handles scaling above one. Scale-to-zero support depends on the Kubernetes release, metrics, and cluster configuration.

How HPA and KEDA differ

The Horizontal Pod Autoscaler (HPA) is a Kubernetes API resource and control-plane controller. It adjusts the replica count of scalable workloads such as Deployments and StatefulSets in response to metrics. Its stable API is autoscaling/v2, which supports resource metrics such as CPU and memory, as well as custom, object, and external metrics when the relevant APIs and providers are available. See the Kubernetes HPA concepts and HPA v2 API reference.

KEDA connects workloads to event sources using scalers and KEDA custom resources. Its operator manages those resources and the HPA lifecycle; its metrics API server exposes scaler metrics for HPA decisions above one replica. In the documented architecture, KEDA provides event-driven activation and deactivation, while HPA ordinarily makes scaling decisions above one replica. KEDA also uses admission webhooks to validate its resources. See KEDA concepts and the KEDA scaler catalog.

Choose based on the signal and scaling range

Decision HPA alone fits when… KEDA fits when…
Demand signal CPU, memory, or a custom, object, or external metric already available to Kubernetes expresses demand. A supported KEDA scaler can read the relevant event source, such as queue activity.
Zero replicas A suitable object or external metric and compatible Kubernetes release and configuration are available. You need event-driven activation from zero and a suitable scaler is available.
Components to operate You want to configure HPA directly and operate the required metrics APIs and adapters. You can operate KEDA’s operator, metrics API server, custom resources, scaler configuration, and any source credentials.
Scaling above one HPA evaluates the configured metrics and behavior policies. KEDA supplies scaler metrics and HPA handles scaling above one replica.

This is a capability-based choice, not a performance ranking. No workload benchmarks are established here.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

When HPA alone is enough

Resource metrics

For CPU and memory, HPA reads resource metrics through metrics.k8s.io, commonly provided by a separately deployed Metrics Server. For CPU utilization targets, it compares usage with requested CPU. If the relevant containers lack CPU requests, utilization can be undefined for that metric, so check resource requests as part of HPA setup.

Custom and external metrics

HPA can also use custom, object, and external metrics. These are not available merely because an HPA manifest names them: the corresponding API must be registered, readable, and backed by a suitable adapter or provider. Confirm the metric path before relying on it to drive replicas.

Multiple metrics and scaling behavior

With autoscaling/v2, an HPA can evaluate multiple metrics and uses the largest replica recommendation it can calculate, subject to the configured maximum. The behavior field allows separate scale-up and scale-down policies, stabilization windows, and tolerance settings. These controls can limit scaling speed and reduce flapping.

The HPA controller runs an intermittent control loop; Kubernetes documentation gives a default synchronization period of 15 seconds. That is a controller default, not a guarantee that a workload will gain ready capacity within 15 seconds: metric availability, scheduling, image startup, and application readiness also affect the time to serve traffic.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Let HPA own the replica count

When HPA manages a workload, omit its declarative spec.replicas field from the workload manifest. Later manifest applies that set a replica count can reset the value HPA has chosen and create unwanted scaling behavior.

When KEDA adds useful capabilities

Event-source scalers

KEDA’s scaler catalog spans messaging, datastores, metrics, data and storage, CI/CD, applications, scheduling, Kubernetes, testing, and monitoring. The catalog labeled KEDA v2.20 is broad, but availability and configuration can vary by scaler and release. Check the documentation matching your deployed KEDA version before choosing a trigger.

Queue consumers and activation from zero

For a queue consumer, KEDA can detect pending work when no worker Pods are running and activate the Deployment. As load grows, it supplies event metrics for HPA to scale the workload above one replica. When the source is idle, suitable configuration can allow the workload to scale back to zero. Workers then pull from the event source; retries and dead-letter handling depend on the application and source, not on autoscaling itself. See KEDA scaling deployments.

CPU and memory triggers are not a zero-wake signal

KEDA’s CPU and memory triggers use the Kubernetes Metrics Server path, but do not support scale-to-zero. With no Pods running, those metrics cannot provide the activation signal. If waking from zero is required, do not rely on a CPU- or memory-only KEDA configuration.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Best Value
Sale
Kubernetes Software - Powerful Container Orchestration Tools T-Shirt
  • Kubernetes is an open platform that automates container orchestration, enabling seamless deployment, automatic scaling, self-healing, and efficient management of applications across servers or clouds with high availability and optimal resource use
  • Kubernetes is perfect for development operations engineers, cloud architects, site reliability engineers, platform engineering teams and infrastructure specialists who build, operate and maintain modern containerized applications in production environments
  • Lightweight, Classic fit, Double-needle sleeve and bottom hem
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Scale-to-zero needs a version and traffic check

CPU and memory alone cannot report demand when no Pods exist. A zero-replica design therefore needs a signal that remains observable without those Pods, such as a suitable object, external, or event metric.

Kubernetes’ HPA concepts documentation describes scale-to-zero through the HPAScaleToZero feature gate and requires at least one object or external metric. A Kubernetes announcement dated September 2, 2026, says HPA scale-to-zero is Beta and enabled by default in Kubernetes v1.37 for suitable object or external metrics. The announcement also specifies that the workload must have been scaled down by HPA; manually setting it to zero leaves it paused. The feature depends on the relevant control-plane components supporting and enabling it. Consult the Kubernetes v1.37 scale-to-zero announcement and verify the actual cluster configuration rather than assuming all Kubernetes versions behave alike.

KEDA’s documented architecture assigns the zero-to-one and one-to-zero transition to the KEDA operator. That is distinct from HPA’s role above one replica, and it still depends on a suitable scaler and release-compatible setup.

Account for the cost of a cold start

Scaling to zero removes idle Pods but introduces a period with no ready application capacity while the metric is observed, a Pod is scheduled, and the application starts. Kubernetes Services do not buffer requests when no Pods are ready. For HTTP or other request-driven workloads that must not lose requests during that interval, provide a separate buffering layer. Queue-backed work may wait durably in its source, depending on that source’s behavior.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A practical decision checklist

  • Start with the demand signal: If CPU, memory, or a metric exposed to Kubernetes is sufficient, HPA may be all you need.
  • Choose KEDA for event-driven scaling: Use it when a supported scaler can represent the workload’s demand more directly, particularly when activation from zero matters.
  • Verify the zero-replica path: Confirm that the signal exists with no Pods, the relevant APIs and providers are available, and the Kubernetes or KEDA configuration supports the transition.
  • Plan for the operating components: HPA may require Metrics Server or metric adapters; KEDA adds its operator, metrics API server, custom resources, scaler setup, and source credentials.
  • Check release-matched documentation: The cited KEDA pages carry different version labels—v2.20 for the scaler catalog, v2.21 for scaling, and v2.22 for concepts. Treat them as separate references and verify compatibility against your deployed versions.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.