Use Kubernetes’ HorizontalPodAutoscaler (HPA) when the job is to change a supported workload’s replica count in response to resource use or a metric exposed through Kubernetes’ metrics APIs. Consider a custom controller when you need durable, domain-specific state reconciliation or lifecycle behavior that HPA’s metric-to-replica interface cannot express. Before building one, check whether the gap is actually a missing metrics adapter, an HPA setting, or a feature unavailable in your cluster’s Kubernetes version.
Start with the change you need Kubernetes to make
The choice is not simply “built-in autoscaling versus more flexibility.” First identify the desired state change, then ask whether HPA can make that change using an available metric and its documented behavior.
| Question | HPA is a strong starting point when… | Consider custom reconciliation when… |
|---|---|---|
| What changes? | You need to adjust replica count on a supported resource with a scale subresource. | The system must coordinate domain-specific objects, sequencing, or lifecycle state beyond replica count. |
| What drives the decision? | CPU, memory, custom, object, or external metrics can represent the load or queue signal. | The policy depends on domain state or transitions that cannot be represented through HPA metrics and behavior. |
| Are built-in controls enough? | Replica bounds, multiple metrics, and scaling behavior meet the requirements. | The required policy remains unexpressible after checking the HPA API and configuration for the cluster’s version. |
| Is the problem in the metrics path? | The required aggregated metrics API or adapter can be installed or corrected. | The controller must reconcile broader desired state, rather than just provide HPA with a scaling signal. |
| What response does the workload need? | Periodic, metric-driven adjustment with HPA’s readiness and stabilization behavior is acceptable. | The application needs repeated reconciliation of domain-specific state or lifecycle operations. |
This is a capability test, not a universal architecture rule: workload latency, safety requirements, and cluster configuration determine whether a documented HPA capability is sufficient.
What HPA can do
Adjust replicas from one or more metrics
HPA is a Kubernetes API resource and control-plane controller that adjusts the desired scale of supported workloads, including Deployments and StatefulSets. The stable autoscaling/v2 API supports resource and custom metrics, and can use multiple metrics. When several metrics are configured, HPA calculates a replica recommendation for each and uses the highest recommendation, subject to the configured replica bounds. Kubernetes: Horizontal Pod Autoscaling
Free tools Windows power users keep installed
One-click scans. No signup required.
#1 Best Overall
Apply more than a direct metric-to-replica formula
For CPU utilization targets, utilization is measured against requested CPU resources, so suitable resource requests matter. HPA also accounts for missing metrics and Pods that are not yet ready. Tolerance around the target and scale-down stabilization can affect the result. The replica count therefore is not necessarily an instantaneous conversion of the latest raw metric sample.
Reconcile periodically, not continuously
The HPA controller checks metrics on a configurable interval through the kube-controller-manager option --horizontal-pod-autoscaler-sync-period. Kubernetes documents a default of 15 seconds. That is the controller’s polling interval, not a guarantee that demand will be served by a ready Pod within 15 seconds; metric collection, scheduling, startup, and configuration also affect end-to-end response. Kubernetes: HPA control loop
Rank #2
Verify the metrics path before writing a controller
A missing or unusable signal does not by itself mean HPA is the wrong tool. Check which Kubernetes metrics API serves the signal and whether it is registered and available:
- Resource metrics: HPA reads resource metrics through
metrics.k8s.io, commonly provided by Metrics Server. - Custom metrics: Custom metrics are served through
custom.metrics.k8s.io, typically by a metrics adapter. - External metrics: External metrics use
external.metrics.k8s.io, also typically provided by an adapter.
Custom and external metric setups depend on API aggregation and API registration. If the signal represents the workload’s demand, investigate whether the corresponding API and adapter can expose it to HPA before adding a separate controller. Kubernetes: HPA metrics APIs
Rank #3
What a custom resource and controller add
A custom resource defines structured objects in the Kubernetes API; it does not perform reconciliation on its own. Paired with a custom controller, it lets users declare desired state while the controller repeatedly works to bring actual Kubernetes objects into line. The Operator pattern applies this model to encode domain knowledge in software. Kubernetes: Operator pattern
That pattern is useful when the desired behavior is genuinely broader than scaling a workload—for example, when a domain-specific API needs to represent application state and the controller must coordinate related resources or lifecycle transitions. This is an architectural recommendation based on the distinction between HPA’s scaling role and controller reconciliation, not a Kubernetes requirement to use an Operator.
Rank #4
Check adjacent scaling options and version-sensitive features
Choose VPA when the desired change is resource sizing
HPA changes replica count. Vertical Pod Autoscaler (VPA) addresses a different goal by adjusting container resource requests and limits. Kubernetes documents VPA as a separately installed component that can use historical utilization, cluster resources, and events. It is not a replacement for HPA when the required action is changing the number of replicas. Kubernetes: Vertical Pod Autoscaling
Confirm scale-to-zero support on the actual cluster
Kubernetes v1.37 documentation, published September 2, 2026, describes HPA scale-to-zero for eligible HPAs using object or external metrics as a beta capability enabled by default. Scaling from zero can involve cold-start delay while HPA observes the metric, Pods are scheduled, and the application starts. Confirm the cluster version and feature state before depending on this behavior. Kubernetes Blog: HPA scale-to-zero in Kubernetes v1.37
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Quick Recap
Best Value
Pre-implementation checklist
- Name the exact desired change: replica count, container resource requests and limits, or broader application lifecycle state.
- Define the signal: identify its owner, units, freshness, and whether it is per-Pod, object, or external; verify that the needed aggregated metrics API is registered.
- Check workload inputs: confirm resource requests when using CPU or memory utilization targets, and consider how Pod readiness affects metrics.
- Validate the policy: compare replica bounds, multiple-metric behavior, tolerance, and scale-up and scale-down stabilization with the workload’s response needs.
- Check cluster capabilities: verify the Kubernetes version and feature state for scale-to-zero or any other required HPA behavior, and account for cold starts where relevant.
- State what HPA cannot express: if the remaining requirement is durable domain state and reconciliation, define a custom resource and controller boundary. If no such requirement remains, avoid adding a new API and controller lifecycle without a demonstrated need.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

