Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To deploy a microservice on Kubernetes, package it as a container image, define its desired state in a Deployment, and give it a stable network endpoint with a Service. Keep environment-specific configuration outside the image, use readiness and liveness probes for different purposes, then verify each rollout before treating it as successful. The example below assumes a stateless HTTP service named orders that listens on port 8080; adapt the image, endpoints, resources, and configuration to your application.

What you need to decide before writing YAML

Kubernetes can keep Pods running and replace them, but it does not decide your service boundaries or application contract. Give each microservice its own image, configuration requirements, health endpoints, resource profile, and ServiceAccount. Build an image once and promote that same immutable image across environments; supply environment-specific values through Kubernetes configuration instead.

  • Service boundary: identify the application process and the network port it serves.
  • Configuration: separate non-confidential settings from passwords, tokens, and keys.
  • Health behavior: define when startup is complete, when a Pod may receive traffic, and what would indicate an unrecoverably stuck process.
  • Operational ownership: decide who manages the cluster, upgrades, certificates, backups, monitoring, and incident response.

The manifests below are a starting point, not universal production values. They assume the application provides /health/startup, /health/ready, and /health/live on port 8080, and that a Secret named orders-db already exists in the same namespace.

Define configuration and workload identity

Use ConfigMaps for non-confidential settings

A ConfigMap stores non-confidential key-value configuration. This example sets a log level outside the image:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
apiVersion: v1
kind: ConfigMap
metadata:
  name: orders-config
  namespace: shop
data:
  LOG_LEVEL: "info"

Keep confidential values in Secrets

Use a Secret for passwords, tokens, keys, and other confidential values. Base64 encoding, which Kubernetes uses for Secret data by default, is not encryption. Secret values are stored unencrypted in etcd unless encryption at rest is configured. Apply encryption at rest, limit access with least-privilege RBAC, and grant Secret access only to workloads that need it. Do not commit manifests containing merely base64-encoded Secret values to source control, and ensure the application does not log a secret after reading it.

For example, the workload below expects a Secret called orders-db with a key named password. Create and manage that Secret using your organization’s approved secret-handling process; the example intentionally does not include a secret value.

Give the workload its own ServiceAccount

A dedicated ServiceAccount gives the workload a distinct Kubernetes identity. The example disables automatic token mounting because the application does not need to call the Kubernetes API. Enable API access only when required, and then grant only the permissions the service needs.

Deploy the microservice and give it a stable endpoint

A Deployment manages ReplicaSets and maintains the desired number of Pods. A Service selects Pods by label and provides clients with a stable endpoint as those Pods are replaced. Keep the Service selector and Deployment Pod labels aligned: a mismatch can leave the Deployment running while the Service has no matching backends.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

This example uses an internal ClusterIP Service. It does not expose the application publicly; add an Ingress, gateway, or load balancer only if the service needs that access. The specific controller and cloud load-balancer behavior depend on your cluster environment.

apiVersion: v1
kind: ServiceAccount
metadata:
  name: orders
  namespace: shop
automountServiceAccountToken: false
---
apiVersion: apps/v1
kind: Deployment
metadata:
  name: orders
  namespace: shop
spec:
  replicas: 3
  selector:
    matchLabels:
      app: orders
  strategy:
    type: RollingUpdate
    rollingUpdate:
      maxUnavailable: 0
      maxSurge: 1
  template:
    metadata:
      labels:
        app: orders
    spec:
      serviceAccountName: orders
      containers:
        - name: orders
          image: registry.example.com/team/orders:1.4.2
          ports:
            - name: http
              containerPort: 8080
          envFrom:
            - configMapRef:
                name: orders-config
          env:
            - name: DB_PASSWORD
              valueFrom:
                secretKeyRef:
                  name: orders-db
                  key: password
          resources:
            requests:
              cpu: 250m
              memory: 256Mi
            limits:
              cpu: "1"
              memory: 512Mi
          startupProbe:
            httpGet:
              path: /health/startup
              port: http
            periodSeconds: 5
            failureThreshold: 30
          readinessProbe:
            httpGet:
              path: /health/ready
              port: http
            periodSeconds: 5
          livenessProbe:
            httpGet:
              path: /health/live
              port: http
            periodSeconds: 10
            failureThreshold: 3
---
apiVersion: v1
kind: Service
metadata:
  name: orders
  namespace: shop
spec:
  type: ClusterIP
  selector:
    app: orders
  ports:
    - name: http
      port: 80
      targetPort: http

The image reference and resource values are illustrative: use an image your cluster can pull and resource requests and limits informed by the service’s behavior. The Deployment’s update settings express a preference not to make an old Pod unavailable before a replacement is ready; they do not guarantee capacity exists to schedule that replacement or that the application will have no errors during an update.

Set probes according to what the application can actually tell Kubernetes

Startup probe: has initialization finished?

Use a startup probe when initialization can take time. While it has not succeeded, Kubernetes does not run the container’s liveness or readiness probes. In the example, a check every five seconds with a failure threshold of 30 allows up to 150 seconds of failed startup checks before Kubernetes treats startup as failed. Tune the values to measured startup behavior rather than copying them blindly.

Readiness probe: should this Pod receive traffic now?

Readiness determines whether a Pod is eligible to receive traffic from matching Service endpoints. A failed readiness check removes the Pod from those endpoints; it does not, by itself, restart the container. Make the endpoint reflect whether the service can handle requests, including during startup and shutdown.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Liveness probe: is the process stuck beyond recovery?

Liveness failure can restart the container. Keep the check cheap and deterministic, and do not make it depend on a flaky downstream service unless restarting this process is the intended recovery action. An overly aggressive liveness check can restart healthy but busy containers and contribute to cascading failures.

Apply the manifests and verify the result

Save the example resources in a file such as orders.yaml, create the namespace and prerequisite configuration, then apply the resources. The image and Secret must be available to the cluster before the Pod can run successfully.

  1. Create the namespace if it does not already exist: kubectl create namespace shop.
  2. Apply the ConfigMap and workload resources: kubectl apply -f orders.yaml.
  3. Wait for the Deployment rollout: kubectl rollout status deployment/orders -n shop --timeout=5m.
  4. Inspect the Deployment and Pods if it does not become ready: kubectl get deployment,pods -n shop.
  5. Check events and container output to diagnose failures: kubectl describe pod -n shop <pod-name> and kubectl logs -n shop <pod-name>.
  6. Confirm that the Service has endpoints: kubectl get endpoints orders -n shop.

Common causes of a failed rollout include an image the cluster cannot pull, a missing Secret or ConfigMap, a probe path or port that the application does not serve, resource requests that cannot be scheduled, and labels that do not match the Service selector. Use Pod events and application logs to distinguish these cases rather than repeatedly restarting the Deployment.

Release updates with a rollback path

For a new release, change the Deployment’s image to the intended immutable tag or digest, then apply the updated manifest. Watch the rollout and check application-level signals such as error rates and request latency before declaring the release healthy. Set a rollback trigger in advance—for example, sustained elevated errors, failed readiness, or an SLO violation—and retain enough rollout history to recover.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
kubectl apply -f orders.yaml
kubectl rollout status deployment/orders -n shop --timeout=5m
kubectl rollout history deployment/orders -n shop

If the new revision is unhealthy, a Deployment rollback can restore the prior revision:

kubectl rollout undo deployment/orders -n shop

A rolling update is a replacement strategy, not a promise of zero downtime or zero errors. Actual availability depends on factors including replica count, readiness behavior, disruption budgets, available cluster capacity, and whether the application tolerates overlapping versions.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Autoscale only after metrics and readiness are usable

A HorizontalPodAutoscaler (HPA) adjusts a scalable workload, such as a Deployment, toward demand. The stable HPA API is autoscaling/v2, which supports resource and other metric sources. For CPU- or memory-based resource scaling, the cluster needs Metrics Server or another Metrics API implementation. Resource metrics support autoscaling and basic inspection; they are not a complete monitoring system.

This example targets average CPU utilization relative to the CPU requests in the Deployment. The replica bounds and utilization target are example choices, not universal recommendations.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: orders
  namespace: shop
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: orders
  minReplicas: 3
  maxReplicas: 12
  metrics:
    - type: Resource
      resource:
        name: cpu
        target:
          type: Utilization
          averageUtilization: 70

Apply the HPA after creating the Deployment, then inspect its current metrics and scaling conditions with kubectl get hpa -n shop and kubectl describe hpa orders -n shop. HPA decisions are affected by Pods that are not yet ready and by missing metrics, which Kubernetes treats conservatively. Startup duration and readiness behavior therefore influence scaling. For systems whose bottleneck is queue depth, request rate, or another application signal, evaluate an appropriate metric source instead of assuming CPU alone represents demand.

Prepare security and operations for production

Restrict access and traffic

  • Protect API traffic with TLS and enforce authentication and authorization.
  • Use least-privilege RBAC, dedicated workload identities, and Pod Security controls.
  • Use NetworkPolicies where appropriate to limit service-to-service traffic.
  • Enable audit logging and decide how long audit records and Kubernetes events must be retained.
  • Set resource requests and limits deliberately so scheduling and resource consumption are managed rather than left implicit.

Plan for availability and recovery

Choose whether the control plane and supporting services are self-managed or provided by a managed Kubernetes service. Before production, assign ownership for node patching, certificate rotation, API availability, etcd or application-data restoration, security advisories, and the observability stack. Plan certificates, API-server load balancing, etcd separation and backups, namespace quotas, DNS capacity, and workload preparation according to the cluster operating model.

Also define backup scope and recovery objectives for application data. A Deployment can recreate Pods, but it does not replace backups of persistent data or a disaster-recovery plan.

Build observability beyond resource metrics

Collect metrics, logs, and traces, and correlate requests across services using request identifiers or equivalent context. Alert on user-facing symptoms as well as infrastructure conditions. CPU and memory metrics alone cannot explain dependency failures, growing queues, or distributed latency.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose an exposure and deployment approach that fits the service

There is no single Kubernetes deployment shape that suits every microservice. Decide these boundaries before standardizing templates:

  • Operational ownership: managed versus self-managed control plane, including who patches and restores it.
  • Exposure: internal Service, gateway, or public load balancer, based on which clients need access.
  • Release method: rolling update, blue/green, or canary, depending on the controls and traffic management available in the environment.
  • Scaling signal: CPU or memory versus application or external metrics.
  • Isolation: namespace, node, network, and identity boundaries appropriate to the workloads.
  • Recovery: backup, rollback, and disaster-recovery objectives for both workloads and data.
  • Observability: whether metrics alone are sufficient or teams need correlated logs and traces to diagnose cross-service behavior.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.