To deploy a microservice on Kubernetes, package it as a container image, define its desired state in a Deployment, and give it a stable network endpoint with a Service. Keep environment-specific configuration outside the image, use readiness and liveness probes for different purposes, then verify each rollout before treating it as successful. The example below assumes a stateless HTTP service named orders that listens on port 8080; adapt the image, endpoints, resources, and configuration to your application.
What you need to decide before writing YAML
Kubernetes can keep Pods running and replace them, but it does not decide your service boundaries or application contract. Give each microservice its own image, configuration requirements, health endpoints, resource profile, and ServiceAccount. Build an image once and promote that same immutable image across environments; supply environment-specific values through Kubernetes configuration instead.
- Service boundary: identify the application process and the network port it serves.
- Configuration: separate non-confidential settings from passwords, tokens, and keys.
- Health behavior: define when startup is complete, when a Pod may receive traffic, and what would indicate an unrecoverably stuck process.
- Operational ownership: decide who manages the cluster, upgrades, certificates, backups, monitoring, and incident response.
The manifests below are a starting point, not universal production values. They assume the application provides /health/startup, /health/ready, and /health/live on port 8080, and that a Secret named orders-db already exists in the same namespace.
Define configuration and workload identity
Use ConfigMaps for non-confidential settings
A ConfigMap stores non-confidential key-value configuration. This example sets a log level outside the image:
#1 Best Overall
apiVersion: v1
kind: ConfigMap
metadata:
name: orders-config
namespace: shop
data:
LOG_LEVEL: "info"
Keep confidential values in Secrets
Use a Secret for passwords, tokens, keys, and other confidential values. Base64 encoding, which Kubernetes uses for Secret data by default, is not encryption. Secret values are stored unencrypted in etcd unless encryption at rest is configured. Apply encryption at rest, limit access with least-privilege RBAC, and grant Secret access only to workloads that need it. Do not commit manifests containing merely base64-encoded Secret values to source control, and ensure the application does not log a secret after reading it.
For example, the workload below expects a Secret called orders-db with a key named password. Create and manage that Secret using your organization’s approved secret-handling process; the example intentionally does not include a secret value.
Give the workload its own ServiceAccount
A dedicated ServiceAccount gives the workload a distinct Kubernetes identity. The example disables automatic token mounting because the application does not need to call the Kubernetes API. Enable API access only when required, and then grant only the permissions the service needs.
Deploy the microservice and give it a stable endpoint
A Deployment manages ReplicaSets and maintains the desired number of Pods. A Service selects Pods by label and provides clients with a stable endpoint as those Pods are replaced. Keep the Service selector and Deployment Pod labels aligned: a mismatch can leave the Deployment running while the Service has no matching backends.
Rank #2
This example uses an internal ClusterIP Service. It does not expose the application publicly; add an Ingress, gateway, or load balancer only if the service needs that access. The specific controller and cloud load-balancer behavior depend on your cluster environment.
apiVersion: v1
kind: ServiceAccount
metadata:
name: orders
namespace: shop
automountServiceAccountToken: false
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: orders
namespace: shop
spec:
replicas: 3
selector:
matchLabels:
app: orders
strategy:
type: RollingUpdate
rollingUpdate:
maxUnavailable: 0
maxSurge: 1
template:
metadata:
labels:
app: orders
spec:
serviceAccountName: orders
containers:
- name: orders
image: registry.example.com/team/orders:1.4.2
ports:
- name: http
containerPort: 8080
envFrom:
- configMapRef:
name: orders-config
env:
- name: DB_PASSWORD
valueFrom:
secretKeyRef:
name: orders-db
key: password
resources:
requests:
cpu: 250m
memory: 256Mi
limits:
cpu: "1"
memory: 512Mi
startupProbe:
httpGet:
path: /health/startup
port: http
periodSeconds: 5
failureThreshold: 30
readinessProbe:
httpGet:
path: /health/ready
port: http
periodSeconds: 5
livenessProbe:
httpGet:
path: /health/live
port: http
periodSeconds: 10
failureThreshold: 3
---
apiVersion: v1
kind: Service
metadata:
name: orders
namespace: shop
spec:
type: ClusterIP
selector:
app: orders
ports:
- name: http
port: 80
targetPort: http
The image reference and resource values are illustrative: use an image your cluster can pull and resource requests and limits informed by the service’s behavior. The Deployment’s update settings express a preference not to make an old Pod unavailable before a replacement is ready; they do not guarantee capacity exists to schedule that replacement or that the application will have no errors during an update.
Set probes according to what the application can actually tell Kubernetes
Startup probe: has initialization finished?
Use a startup probe when initialization can take time. While it has not succeeded, Kubernetes does not run the container’s liveness or readiness probes. In the example, a check every five seconds with a failure threshold of 30 allows up to 150 seconds of failed startup checks before Kubernetes treats startup as failed. Tune the values to measured startup behavior rather than copying them blindly.
Readiness probe: should this Pod receive traffic now?
Readiness determines whether a Pod is eligible to receive traffic from matching Service endpoints. A failed readiness check removes the Pod from those endpoints; it does not, by itself, restart the container. Make the endpoint reflect whether the service can handle requests, including during startup and shutdown.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchPC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Rank #3
Liveness probe: is the process stuck beyond recovery?
Liveness failure can restart the container. Keep the check cheap and deterministic, and do not make it depend on a flaky downstream service unless restarting this process is the intended recovery action. An overly aggressive liveness check can restart healthy but busy containers and contribute to cascading failures.
Apply the manifests and verify the result
Save the example resources in a file such as orders.yaml, create the namespace and prerequisite configuration, then apply the resources. The image and Secret must be available to the cluster before the Pod can run successfully.
- Create the namespace if it does not already exist:
kubectl create namespace shop. - Apply the ConfigMap and workload resources:
kubectl apply -f orders.yaml. - Wait for the Deployment rollout:
kubectl rollout status deployment/orders -n shop --timeout=5m. - Inspect the Deployment and Pods if it does not become ready:
kubectl get deployment,pods -n shop. - Check events and container output to diagnose failures:
kubectl describe pod -n shop <pod-name>andkubectl logs -n shop <pod-name>. - Confirm that the Service has endpoints:
kubectl get endpoints orders -n shop.
Common causes of a failed rollout include an image the cluster cannot pull, a missing Secret or ConfigMap, a probe path or port that the application does not serve, resource requests that cannot be scheduled, and labels that do not match the Service selector. Use Pod events and application logs to distinguish these cases rather than repeatedly restarting the Deployment.
Release updates with a rollback path
For a new release, change the Deployment’s image to the intended immutable tag or digest, then apply the updated manifest. Watch the rollout and check application-level signals such as error rates and request latency before declaring the release healthy. Set a rollback trigger in advance—for example, sustained elevated errors, failed readiness, or an SLO violation—and retain enough rollout history to recover.
Free tools Windows power users keep installed
One-click scans. No signup required.
Rank #4
kubectl apply -f orders.yaml
kubectl rollout status deployment/orders -n shop --timeout=5m
kubectl rollout history deployment/orders -n shop
If the new revision is unhealthy, a Deployment rollback can restore the prior revision:
kubectl rollout undo deployment/orders -n shop
A rolling update is a replacement strategy, not a promise of zero downtime or zero errors. Actual availability depends on factors including replica count, readiness behavior, disruption budgets, available cluster capacity, and whether the application tolerates overlapping versions.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Autoscale only after metrics and readiness are usable
A HorizontalPodAutoscaler (HPA) adjusts a scalable workload, such as a Deployment, toward demand. The stable HPA API is autoscaling/v2, which supports resource and other metric sources. For CPU- or memory-based resource scaling, the cluster needs Metrics Server or another Metrics API implementation. Resource metrics support autoscaling and basic inspection; they are not a complete monitoring system.
This example targets average CPU utilization relative to the CPU requests in the Deployment. The replica bounds and utilization target are example choices, not universal recommendations.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Best Value
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: orders
namespace: shop
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: orders
minReplicas: 3
maxReplicas: 12
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 70
Apply the HPA after creating the Deployment, then inspect its current metrics and scaling conditions with kubectl get hpa -n shop and kubectl describe hpa orders -n shop. HPA decisions are affected by Pods that are not yet ready and by missing metrics, which Kubernetes treats conservatively. Startup duration and readiness behavior therefore influence scaling. For systems whose bottleneck is queue depth, request rate, or another application signal, evaluate an appropriate metric source instead of assuming CPU alone represents demand.
Prepare security and operations for production
Restrict access and traffic
- Protect API traffic with TLS and enforce authentication and authorization.
- Use least-privilege RBAC, dedicated workload identities, and Pod Security controls.
- Use NetworkPolicies where appropriate to limit service-to-service traffic.
- Enable audit logging and decide how long audit records and Kubernetes events must be retained.
- Set resource requests and limits deliberately so scheduling and resource consumption are managed rather than left implicit.
Plan for availability and recovery
Choose whether the control plane and supporting services are self-managed or provided by a managed Kubernetes service. Before production, assign ownership for node patching, certificate rotation, API availability, etcd or application-data restoration, security advisories, and the observability stack. Plan certificates, API-server load balancing, etcd separation and backups, namespace quotas, DNS capacity, and workload preparation according to the cluster operating model.
Also define backup scope and recovery objectives for application data. A Deployment can recreate Pods, but it does not replace backups of persistent data or a disaster-recovery plan.
Build observability beyond resource metrics
Collect metrics, logs, and traces, and correlate requests across services using request identifiers or equivalent context. Alert on user-facing symptoms as well as infrastructure conditions. CPU and memory metrics alone cannot explain dependency failures, growing queues, or distributed latency.
Choose an exposure and deployment approach that fits the service
There is no single Kubernetes deployment shape that suits every microservice. Decide these boundaries before standardizing templates:
Quick Recap
- Operational ownership: managed versus self-managed control plane, including who patches and restores it.
- Exposure: internal Service, gateway, or public load balancer, based on which clients need access.
- Release method: rolling update, blue/green, or canary, depending on the controls and traffic management available in the environment.
- Scaling signal: CPU or memory versus application or external metrics.
- Isolation: namespace, node, network, and identity boundaries appropriate to the workloads.
- Recovery: backup, rollback, and disaster-recovery objectives for both workloads and data.
- Observability: whether metrics alone are sufficient or teams need correlated logs and traces to diagnose cross-service behavior.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

