Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To deploy a scalable Go application on Kubernetes, package it as an immutable container image, run it in a Deployment, expose it through a Service, and use a HorizontalPodAutoscaler (HPA) to adjust the number of Pods as demand changes. For resource-based HPA scaling, every relevant container needs a CPU or memory request, and the cluster needs Metrics Server or an equivalent resource metrics API. Tune requests, limits, replica targets, probes, and rollout settings using representative load tests and production telemetry; there is no universal CPU or memory setting for a Go container.

How the scaling pieces fit together

A scalable deployment involves several distinct control points. The HPA changes the number of Pods for a workload such as a Deployment. A Service provides a stable endpoint for reaching the Pods that match its selector. Node autoscaling adds or removes cluster capacity when Pods cannot be scheduled or nodes are underused. Vertical Pod Autoscaling (VPA) addresses per-Pod resource sizing rather than replica count.

  • Deployment: declares the desired Pod template and manages replicated Pods and their updates.
  • Service: selects matching Pods and gives clients a stable in-cluster endpoint as Pods are replaced or scaled.
  • HPA: periodically adjusts a workload’s replica count based on configured resource, custom, or external metrics.
  • VPA: adjusts resource requests and, depending on its configuration, may require Pods to be recreated to apply changes.
  • Node autoscaling: changes the amount of worker-node capacity; it does not directly change the Deployment’s replica count.

These mechanisms can complement one another, but they do not substitute for one another. Adding replicas cannot help if the cluster has no capacity to schedule them, and adding nodes alone does not create more application Pods.

Prepare the Go service and container image

Build the application as a small, stateless HTTP or gRPC service where practical. Keep state that must survive Pod replacement in an appropriate external system rather than in a container’s local filesystem. Publish an immutable image reference, such as a versioned tag or image digest, so a Deployment rollout uses a known build rather than a mutable tag that may change between pulls.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Decide which port the process listens on and make configuration available through environment variables, ConfigMaps, or Secrets as appropriate. Do not bake credentials into the image. The application should handle termination gracefully: stop accepting new work, complete or safely abandon in-flight requests within the termination window, and exit when Kubernetes stops the container.

Create a Deployment with measured resource settings

Use a Deployment to declare the Pod template, labels, initial replica count, container image, port, configuration, probes, and resource requests and limits. The following is a template, not a ready-to-apply manifest: replace every REPLACE_... value with a valid value for your cluster and application. The numeric request and limit values cannot be chosen reliably from the language alone; establish them through representative load tests and production telemetry.

apiVersion: apps/v1
kind: Deployment
metadata:
  name: go-api
spec:
  replicas: 2
  selector:
    matchLabels:
      app: go-api
  strategy:
    type: RollingUpdate
    rollingUpdate:
      maxUnavailable: 0
      maxSurge: 1
  template:
    metadata:
      labels:
        app: go-api
    spec:
      terminationGracePeriodSeconds: 30
      containers:
        - name: app
          image: REPLACE_WITH_IMMUTABLE_IMAGE_REFERENCE
          ports:
            - name: http
              containerPort: REPLACE_WITH_LISTEN_PORT
          envFrom:
            - configMapRef:
                name: go-api-config
          resources:
            requests:
              cpu: REPLACE_WITH_MEASURED_CPU_REQUEST
              memory: REPLACE_WITH_MEASURED_MEMORY_REQUEST
            limits:
              cpu: REPLACE_WITH_VALIDATED_CPU_LIMIT
              memory: REPLACE_WITH_VALIDATED_MEMORY_LIMIT
          startupProbe:
            httpGet:
              path: /startup
              port: http
            periodSeconds: 5
            failureThreshold: 30
          readinessProbe:
            httpGet:
              path: /ready
              port: http
            periodSeconds: 5
          livenessProbe:
            httpGet:
              path: /live
              port: http
            periodSeconds: 10

The probe paths are examples of application endpoints, not built-in Go or Kubernetes endpoints. Implement them, or change the probe type and settings to match the service. A startup probe gives a slow-starting process time to initialize before liveness checks begin. Readiness should remain false until the instance can serve requests, including any required warm-up or dependency checks; an unready Pod is not considered ready to receive Service traffic. Liveness should detect a process that needs restarting, not merely a temporarily unavailable downstream dependency.

The rollout settings above illustrate a conservative availability choice, not a universal setting. A rollout that keeps an old Pod available while creating a replacement needs enough cluster capacity and must be compatible with the application’s shutdown behavior. Align rollout policy with replica count, capacity, and any Pod disruption budget.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Expose the Pods through a Service

A Service selector must match the labels on the Deployment’s Pod template. For example:

apiVersion: v1
kind: Service
metadata:
  name: go-api
spec:
  selector:
    app: go-api
  ports:
    - name: http
      port: 80
      targetPort: http
  type: ClusterIP

This creates an in-cluster endpoint; it does not by itself publish the application outside the cluster. Add an Ingress or Gateway only if external routing is required, and configure the relevant controller and network policy for the cluster.

Enable and configure horizontal Pod autoscaling

Before creating a resource-based HPA, ensure Metrics Server or another compatible resource metrics API is available. The Kubernetes Metrics Server collects resource metrics from kubelets and exposes them through the Kubernetes API. Check that metrics are actually available, for example with kubectl top pods; an installed component that is not returning metrics is not enough. The HPA controller is a periodic control loop, with a 15-second default sync period in current Kubernetes HPA documentation, so scaling is not instantaneous.

Here is an autoscaling/v2 CPU HPA template. The 70% target is only an example of manifest syntax, not a Go-specific recommendation or a target that is appropriate for every workload. Choose a target and replica bounds based on load testing, startup time, latency objectives, and available capacity.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: go-api
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: go-api
  minReplicas: 2
  maxReplicas: 12
  metrics:
    - type: Resource
      resource:
        name: cpu
        target:
          type: Utilization
          averageUtilization: 70

For a utilization target, Kubernetes calculates utilization relative to the corresponding resource request. Each relevant container must have a request for the resource used by the metric; if a container lacks that request, Kubernetes cannot define the Pod’s utilization for that metric. Set requests for sidecars as well as the application container when those containers are included in the workload’s resource accounting. CPU or memory utilization targets also need their corresponding metrics available.

If the workload should scale on queue depth, request rate, or latency instead, configure a custom or external metric and provide the corresponding metrics API or adapter. Those signals are not supplied simply by installing Metrics Server. Choose a signal that represents actionable demand and test its behavior under bursts, delayed processing, and metric gaps.

Once the HPA controls replica count, do not continuously apply a Deployment manifest with a fixed spec.replicas: each apply can fight the HPA and cause replica-count thrashing. Omit that field from the repeatedly applied Deployment manifest after autoscaling is enabled, while retaining the HPA’s minimum and maximum replica settings as the scaling bounds.

Choose requests, limits, and targets from evidence

Resource requests and limits have different roles. Requests inform scheduling and are the denominator for HPA utilization targets; limits cap container consumption for the resources where limits are enforced. Kubernetes documentation notes that setting Pod requests and limits appropriately helps both HPA and Cluster Autoscaler make better decisions. There is no generally valid CPU or memory request for a Go service: a small idle process may behave very differently from one under representative concurrency, payload sizes, garbage-collection activity, and dependency latency.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  1. Load-test a build with representative request mix, concurrency, and dependency behavior.
  2. Measure per-container CPU and memory at steady state and during peaks, startup, and recovery. Include sidecars when present.
  3. Set requests to reflect the capacity the scheduler should reserve and the baseline from which utilization scaling is calculated.
  4. Set limits only after validating the effect of throttling and memory enforcement on latency, throughput, and restart behavior.
  5. Observe production telemetry and revise the settings when workload shape, traffic, or application behavior changes.

Do not treat an HPA target as a direct promise about latency or throughput. The control loop reacts to metrics, and new Pods take time to schedule, start, pass readiness, and serve traffic. Set capacity headroom and stabilization behavior for the application’s burst pattern rather than expecting replicas to appear at the first instant of a spike.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Distinguish HPA, VPA, and node autoscaling

Approach What it changes Signals and prerequisites Reaction and disruption considerations Cost and capacity role
Manual replica changes Number of workload Pods. An operator or deployment process changes the replica count; no autoscaling metrics API is needed. Change occurs when applied, but people or automation must respond to demand. Normal Pod scheduling and startup still apply. Simple, but cannot respond automatically to changing load. Pods still need schedulable cluster capacity.
HPA Number of workload Pods. Resource metrics require Metrics Server or a compatible resource metrics API; custom or external signals need their own APIs and adapters. Periodic control loop; target choice, metric behavior, stabilization settings, and Pod startup affect response. Scaling out does not guarantee immediate ready capacity. Matches replica count to configured demand signals, but needs node capacity and appropriate resource requests.
VPA Resource requests, and potentially limits depending on configuration, for individual Pods. Uses observed resource behavior; requires a VPA implementation and its configured recommendations or updates. Applying new resource settings may require Pod recreation, depending on update mode. Account for disruption and workload availability. Can improve per-Pod sizing; it does not add replicas or worker nodes. Stable since Kubernetes v1.25.
Node autoscaling Worker-node capacity. Requires a cluster autoscaling component and a supported infrastructure or cloud-provider integration. Provisioning nodes takes time, and placement is subject to scheduling constraints, quotas, and availability-zone capacity. Provides capacity for Pods that cannot otherwise be scheduled; does not itself scale the application workload.

Container-resource metrics became stable in Kubernetes v1.30, and VPA became stable in Kubernetes v1.25, according to current Kubernetes documentation. Availability of particular autoscaling components and integrations depends on the Kubernetes distribution and environment.

Troubleshoot an HPA that does not scale

Inspect the HPA status and events first, then follow the dependency chain from metrics to scheduling. Useful commands include:

kubectl get hpa go-api
kubectl describe hpa go-api
kubectl get deployment go-api
kubectl get pods -l app=go-api
kubectl top pods
kubectl describe pods -l app=go-api
  • HPA shows unknown or missing metrics: confirm Metrics Server or the intended metrics adapter is healthy and the relevant API is returning data. For a custom or external metric, verify its API and adapter rather than assuming resource metrics provide it.
  • CPU or memory utilization cannot be calculated: verify that every relevant container has a request for the resource targeted by the HPA.
  • Replicas do not increase despite load: inspect HPA events, target and metric values, minimum and maximum replicas, and whether a repeatedly applied Deployment manifest is resetting spec.replicas.
  • Desired replicas rise but Pods stay pending: inspect Pod scheduling events, quotas, node capacity, affinity and topology constraints, and availability-zone capacity. Configure node autoscaling if additional nodes are needed, and check its interaction with disruption budgets and quotas.
  • Replicas rise but service capacity does not: inspect startup and readiness failures, application warm-up, dependency health, and whether the Service selector matches the Pod labels.

Autoscaling is only useful when the entire path works: a metric must represent demand, the HPA must be able to act on it, Pods must become ready, the Service must select them, and the cluster must have capacity to run them.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.