Recommended Free Tools
To deploy a scalable Go application on Kubernetes, package it as an immutable container image, run it in a Deployment, expose it through a Service, and use a HorizontalPodAutoscaler (HPA) to adjust the number of Pods as demand changes. For resource-based HPA scaling, every relevant container needs a CPU or memory request, and the cluster needs Metrics Server or an equivalent resource metrics API. Tune requests, limits, replica targets, probes, and rollout settings using representative load tests and production telemetry; there is no universal CPU or memory setting for a Go container.
How the scaling pieces fit together
A scalable deployment involves several distinct control points. The HPA changes the number of Pods for a workload such as a Deployment. A Service provides a stable endpoint for reaching the Pods that match its selector. Node autoscaling adds or removes cluster capacity when Pods cannot be scheduled or nodes are underused. Vertical Pod Autoscaling (VPA) addresses per-Pod resource sizing rather than replica count.
- Deployment: declares the desired Pod template and manages replicated Pods and their updates.
- Service: selects matching Pods and gives clients a stable in-cluster endpoint as Pods are replaced or scaled.
- HPA: periodically adjusts a workload’s replica count based on configured resource, custom, or external metrics.
- VPA: adjusts resource requests and, depending on its configuration, may require Pods to be recreated to apply changes.
- Node autoscaling: changes the amount of worker-node capacity; it does not directly change the Deployment’s replica count.
These mechanisms can complement one another, but they do not substitute for one another. Adding replicas cannot help if the cluster has no capacity to schedule them, and adding nodes alone does not create more application Pods.
Prepare the Go service and container image
Build the application as a small, stateless HTTP or gRPC service where practical. Keep state that must survive Pod replacement in an appropriate external system rather than in a container’s local filesystem. Publish an immutable image reference, such as a versioned tag or image digest, so a Deployment rollout uses a known build rather than a mutable tag that may change between pulls.
#1 Best Overall
Decide which port the process listens on and make configuration available through environment variables, ConfigMaps, or Secrets as appropriate. Do not bake credentials into the image. The application should handle termination gracefully: stop accepting new work, complete or safely abandon in-flight requests within the termination window, and exit when Kubernetes stops the container.
Create a Deployment with measured resource settings
Use a Deployment to declare the Pod template, labels, initial replica count, container image, port, configuration, probes, and resource requests and limits. The following is a template, not a ready-to-apply manifest: replace every REPLACE_... value with a valid value for your cluster and application. The numeric request and limit values cannot be chosen reliably from the language alone; establish them through representative load tests and production telemetry.
apiVersion: apps/v1
kind: Deployment
metadata:
name: go-api
spec:
replicas: 2
selector:
matchLabels:
app: go-api
strategy:
type: RollingUpdate
rollingUpdate:
maxUnavailable: 0
maxSurge: 1
template:
metadata:
labels:
app: go-api
spec:
terminationGracePeriodSeconds: 30
containers:
- name: app
image: REPLACE_WITH_IMMUTABLE_IMAGE_REFERENCE
ports:
- name: http
containerPort: REPLACE_WITH_LISTEN_PORT
envFrom:
- configMapRef:
name: go-api-config
resources:
requests:
cpu: REPLACE_WITH_MEASURED_CPU_REQUEST
memory: REPLACE_WITH_MEASURED_MEMORY_REQUEST
limits:
cpu: REPLACE_WITH_VALIDATED_CPU_LIMIT
memory: REPLACE_WITH_VALIDATED_MEMORY_LIMIT
startupProbe:
httpGet:
path: /startup
port: http
periodSeconds: 5
failureThreshold: 30
readinessProbe:
httpGet:
path: /ready
port: http
periodSeconds: 5
livenessProbe:
httpGet:
path: /live
port: http
periodSeconds: 10
The probe paths are examples of application endpoints, not built-in Go or Kubernetes endpoints. Implement them, or change the probe type and settings to match the service. A startup probe gives a slow-starting process time to initialize before liveness checks begin. Readiness should remain false until the instance can serve requests, including any required warm-up or dependency checks; an unready Pod is not considered ready to receive Service traffic. Liveness should detect a process that needs restarting, not merely a temporarily unavailable downstream dependency.
The rollout settings above illustrate a conservative availability choice, not a universal setting. A rollout that keeps an old Pod available while creating a replacement needs enough cluster capacity and must be compatible with the application’s shutdown behavior. Align rollout policy with replica count, capacity, and any Pod disruption budget.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Clear out junk files and repair common Windows errors3Fix the driver behind crashes, sound loss and screen glitchesExpose the Pods through a Service
A Service selector must match the labels on the Deployment’s Pod template. For example:
apiVersion: v1
kind: Service
metadata:
name: go-api
spec:
selector:
app: go-api
ports:
- name: http
port: 80
targetPort: http
type: ClusterIP
This creates an in-cluster endpoint; it does not by itself publish the application outside the cluster. Add an Ingress or Gateway only if external routing is required, and configure the relevant controller and network policy for the cluster.
Rank #3
Enable and configure horizontal Pod autoscaling
Before creating a resource-based HPA, ensure Metrics Server or another compatible resource metrics API is available. The Kubernetes Metrics Server collects resource metrics from kubelets and exposes them through the Kubernetes API. Check that metrics are actually available, for example with kubectl top pods; an installed component that is not returning metrics is not enough. The HPA controller is a periodic control loop, with a 15-second default sync period in current Kubernetes HPA documentation, so scaling is not instantaneous.
Here is an autoscaling/v2 CPU HPA template. The 70% target is only an example of manifest syntax, not a Go-specific recommendation or a target that is appropriate for every workload. Choose a target and replica bounds based on load testing, startup time, latency objectives, and available capacity.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: go-api
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: go-api
minReplicas: 2
maxReplicas: 12
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 70
For a utilization target, Kubernetes calculates utilization relative to the corresponding resource request. Each relevant container must have a request for the resource used by the metric; if a container lacks that request, Kubernetes cannot define the Pod’s utilization for that metric. Set requests for sidecars as well as the application container when those containers are included in the workload’s resource accounting. CPU or memory utilization targets also need their corresponding metrics available.
If the workload should scale on queue depth, request rate, or latency instead, configure a custom or external metric and provide the corresponding metrics API or adapter. Those signals are not supplied simply by installing Metrics Server. Choose a signal that represents actionable demand and test its behavior under bursts, delayed processing, and metric gaps.
Once the HPA controls replica count, do not continuously apply a Deployment manifest with a fixed spec.replicas: each apply can fight the HPA and cause replica-count thrashing. Omit that field from the repeatedly applied Deployment manifest after autoscaling is enabled, while retaining the HPA’s minimum and maximum replica settings as the scaling bounds.
Choose requests, limits, and targets from evidence
Resource requests and limits have different roles. Requests inform scheduling and are the denominator for HPA utilization targets; limits cap container consumption for the resources where limits are enforced. Kubernetes documentation notes that setting Pod requests and limits appropriately helps both HPA and Cluster Autoscaler make better decisions. There is no generally valid CPU or memory request for a Go service: a small idle process may behave very differently from one under representative concurrency, payload sizes, garbage-collection activity, and dependency latency.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Best Value
- Load-test a build with representative request mix, concurrency, and dependency behavior.
- Measure per-container CPU and memory at steady state and during peaks, startup, and recovery. Include sidecars when present.
- Set requests to reflect the capacity the scheduler should reserve and the baseline from which utilization scaling is calculated.
- Set limits only after validating the effect of throttling and memory enforcement on latency, throughput, and restart behavior.
- Observe production telemetry and revise the settings when workload shape, traffic, or application behavior changes.
Do not treat an HPA target as a direct promise about latency or throughput. The control loop reacts to metrics, and new Pods take time to schedule, start, pass readiness, and serve traffic. Set capacity headroom and stabilization behavior for the application’s burst pattern rather than expecting replicas to appear at the first instant of a spike.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Distinguish HPA, VPA, and node autoscaling
| Approach | What it changes | Signals and prerequisites | Reaction and disruption considerations | Cost and capacity role |
|---|---|---|---|---|
| Manual replica changes | Number of workload Pods. | An operator or deployment process changes the replica count; no autoscaling metrics API is needed. | Change occurs when applied, but people or automation must respond to demand. Normal Pod scheduling and startup still apply. | Simple, but cannot respond automatically to changing load. Pods still need schedulable cluster capacity. |
| HPA | Number of workload Pods. | Resource metrics require Metrics Server or a compatible resource metrics API; custom or external signals need their own APIs and adapters. | Periodic control loop; target choice, metric behavior, stabilization settings, and Pod startup affect response. Scaling out does not guarantee immediate ready capacity. | Matches replica count to configured demand signals, but needs node capacity and appropriate resource requests. |
| VPA | Resource requests, and potentially limits depending on configuration, for individual Pods. | Uses observed resource behavior; requires a VPA implementation and its configured recommendations or updates. | Applying new resource settings may require Pod recreation, depending on update mode. Account for disruption and workload availability. | Can improve per-Pod sizing; it does not add replicas or worker nodes. Stable since Kubernetes v1.25. |
| Node autoscaling | Worker-node capacity. | Requires a cluster autoscaling component and a supported infrastructure or cloud-provider integration. | Provisioning nodes takes time, and placement is subject to scheduling constraints, quotas, and availability-zone capacity. | Provides capacity for Pods that cannot otherwise be scheduled; does not itself scale the application workload. |
Container-resource metrics became stable in Kubernetes v1.30, and VPA became stable in Kubernetes v1.25, according to current Kubernetes documentation. Availability of particular autoscaling components and integrations depends on the Kubernetes distribution and environment.
Troubleshoot an HPA that does not scale
Inspect the HPA status and events first, then follow the dependency chain from metrics to scheduling. Useful commands include:
kubectl get hpa go-api
kubectl describe hpa go-api
kubectl get deployment go-api
kubectl get pods -l app=go-api
kubectl top pods
kubectl describe pods -l app=go-api
- HPA shows unknown or missing metrics: confirm Metrics Server or the intended metrics adapter is healthy and the relevant API is returning data. For a custom or external metric, verify its API and adapter rather than assuming resource metrics provide it.
- CPU or memory utilization cannot be calculated: verify that every relevant container has a request for the resource targeted by the HPA.
- Replicas do not increase despite load: inspect HPA events, target and metric values, minimum and maximum replicas, and whether a repeatedly applied Deployment manifest is resetting
spec.replicas. - Desired replicas rise but Pods stay pending: inspect Pod scheduling events, quotas, node capacity, affinity and topology constraints, and availability-zone capacity. Configure node autoscaling if additional nodes are needed, and check its interaction with disruption budgets and quotas.
- Replicas rise but service capacity does not: inspect startup and readiness failures, application warm-up, dependency health, and whether the Service selector matches the Pod labels.
Autoscaling is only useful when the entire path works: a metric must represent demand, the HPA must be able to act on it, Pods must become ready, the Service must select them, and the cluster must have capacity to run them.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchQuick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

