To run vLLM on Kubernetes, start with a cluster that has allocatable GPUs, persistent or high-throughput storage for model files, and a Kubernetes Deployment and Service exposing vLLM’s OpenAI-compatible API on port 8000. Add Helm when you need repeatable releases; choose the vLLM production stack when you also need model-aware routing, dashboards, or KV-cache offloading. For inference distributed across nodes, use a KubeRay-backed RayCluster or LeaderWorkerSet (LWS). Scaling reliably requires coordinating application replicas, GPU capacity, and—where used—Ray autoscaling.
What you need before deploying vLLM
- A GPU-capable Kubernetes cluster: Install the NVIDIA Kubernetes Device Plugin and confirm that GPU resources are allocatable to workloads. An NVIDIA GPU for vLLM inference must have enough VRAM for the chosen model and its workload; model size, context length, batching, and parallelism affect the fit. Interconnect topology matters when inference spans multiple GPUs or nodes.
- Model access and storage: Plan where model weights and cache will live. A PersistentVolumeClaim is optional in the basic vLLM deployment example, but persistent or high-throughput storage can avoid repeatedly downloading weights when pods restart. Gated Hugging Face models require a token; store it in a Kubernetes Secret rather than embedding it in a manifest.
- Resource and network settings: Set CPU, memory, GPU, shared-memory, and ephemeral-storage requests deliberately. Decide how clients will reach the API, and keep it private until readiness checks and access controls are in place.
There is no generally applicable GPU count or performance figure for vLLM on Kubernetes. The required hardware depends on the model, available VRAM, context length, batching, parallelism, and traffic pattern. Benchmark the workload you intend to serve rather than sizing from a universal tokens-per-second or cost estimate.
Choose a deployment approach
| Approach | Best fit | What it provides | Main trade-off |
|---|---|---|---|
| Native Kubernetes manifests | A clear, inspectable starting point for a single-node vLLM server | A Deployment, Service, GPU requests, optional model-cache volume, and access to the API | You own packaging, configuration reuse, and operational add-ons. |
| Helm chart | Teams that want repeatable releases and environment-specific configuration | Packaged deployment with values overrides; the documented example includes health probes and defaults to one replica with one requested nvidia.com/gpu. |
You must manage chart and image versions and understand the chart’s values. |
| vLLM production stack | Teams seeking a reference architecture with routing and observability features | Helm-based deployment, Grafana dashboards, multimodel support, model- and prefix-aware routing, fast bootstrapping, and optional LMCache KV-cache offloading | More components and settings to operate than a basic Deployment. |
| KubeRay / RayCluster or LWS | Distributed serving when a workload needs multiple nodes | Patterns for multi-node inference and parallel execution across GPUs and nodes | Requires deliberate GPU-topology, networking, parallelism, and autoscaling design. |
The vLLM production stack is an officially released, production-optimized codebase under the vLLM project. It wraps upstream vLLM without modifying its code. Native manifests remain useful when you want the fewest moving parts; Helm packages configuration for reuse; the production stack adds routing and operational features. These options are not interchangeable in complexity or scope.
Deploy a basic vLLM service with Kubernetes
The official vLLM Kubernetes guide uses a Deployment and Service as its basic pattern. Its example uses the vllm/vllm-openai:latest image and mistralai/Mistral-7B-Instruct-v0.3 as a model example. Treat those as documentation examples, not production version recommendations: pin a tested image version for production.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →#1 Best Overall
- Verify GPU capacity. Confirm the NVIDIA device plugin is working and the nodes expose allocatable GPU resources. The pod’s GPU request must match what the cluster can schedule.
- Prepare model access and cache. Configure the model identifier and, for gated Hugging Face models, mount credentials from a Kubernetes Secret. Add a PersistentVolumeClaim if you want model cache to survive pod replacement or avoid repeated downloads.
- Create the Deployment. Set the vLLM container image, model, GPU resource request, and CPU and memory requests. The guide mounts a
/dev/shmvolume for tensor-parallel inference; configure shared memory when the selected inference setup requires it. - Expose the API with a Service. Target the vLLM container on port
8000. Keep the Service internal until you have verified health and configured appropriate access controls. - Wait for readiness and test a request. The guide’s startup output includes “Application startup complete.” Confirm the pod is ready, then send a request to the OpenAI-compatible
/v1/completionsendpoint through the Service before opening external ingress.
The official guide describes Kubernetes as a way to scale and manage GPU-backed vLLM models. A Deployment can manage pod replicas, but increasing replicas is useful only when the cluster has GPU capacity and each replica can load and serve the model.
When Helm or the production stack is a better fit
Use Helm for repeatable releases
Helm is a Kubernetes package manager that automates application deployment. The official vLLM Helm documentation requires a running cluster, the NVIDIA Kubernetes Device Plugin, available GPU resources, and model storage. Its example chart has /health liveness and readiness probes, defaults to one replica, and requests one nvidia.com/gpu. Those are example defaults, not universal requirements.
Rank #2
Use chart values to keep environment-specific settings manageable, and pin both chart and image versions in production. Review probe behavior, resource requests, model-cache handling, and Secret configuration for the chart version you deploy.
Use the production stack for routing and operational features
The vLLM production stack is a stronger fit when you need to serve multiple models or want routing and dashboards included in a reference setup. Its documented features include model-aware and prefix-aware routing, fast bootstrapping, Grafana dashboards, and optional KV-cache offloading through LMCache. The documented installation uses the vLLM Helm repository and the vllm/vllm-stack chart. Check chart and image versions and review security settings before rollout.
Scale beyond one node with KubeRay or LeaderWorkerSet
KubeRay and RayCluster
The production-stack Helm reference can set raySpec.enabled: true to deploy a model as a multi-node RayCluster through KubeRay rather than as a standard Deployment. Its values include requested GPU count and type, shared-memory size, tensor-parallel size, maximum model length, maximum sequences, prefix caching, chunked prefill, and GPU memory utilization. These settings should reflect the model and hardware topology; they are not safe to copy blindly between workloads.
LeaderWorkerSet
LeaderWorkerSet is another Kubernetes-native pattern for distributed inference. The vLLM LWS guide presents multi-host inference as a major use case. One example requires at least two nodes, each with eight GPUs, and configures tensor parallelism of 8 with pipeline parallelism of 2 for a large model. Those figures describe that example configuration only; they are not a general minimum for LWS or vLLM.
Rank #4
Use multi-node execution when a model or target workload cannot be served suitably on one node. Before deploying, match tensor and pipeline parallelism to model memory needs and GPU interconnects, and verify that the required node and GPU topology can be scheduled together.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Design autoscaling and operations together
Scaling vLLM is not a single autoscaler setting. Ray’s Kubernetes production guidance distinguishes Serve application autoscaling from cluster provisioning and discusses the relationship between Ray autoscaling and the Kubernetes Cluster Autoscaler. Application replicas can respond to demand only if Ray and Kubernetes can supply the needed GPU nodes; node provisioning and model startup also add delay.
- Watch demand: Track queueing and latency alongside request volume so you can see pressure before it becomes sustained failure.
- Watch capacity: Monitor replica counts, GPU utilization, GPU memory and KV-cache pressure, as well as node and GPU provisioning time.
- Watch service health: Track readiness, error rates, and the time from pod creation to a model being able to serve requests.
- Coordinate scaling layers: Set application, Ray, and Kubernetes scaling behavior so that added replicas have GPU capacity, while accounting for model download and pod-start latency.
- Protect and stabilize deployments: Pin versions, restrict network access, use Secrets for gated-model credentials, and test health probes and endpoint behavior before exposing the service.
The official deployment guidance does not establish a universal throughput, latency, utilization, or cost figure. Measure with the actual model, GPU type, context lengths, batching, parallelism, and traffic shape you expect to serve.
Quick Recap
Practical selection rule
- Start with a native Deployment and Service when you need a simple, inspectable single-node service.
- Choose Helm when repeatable packaging and versioned, environment-specific configuration are the main need.
- Choose the vLLM production stack when its routing, dashboards, multimodel support, or optional KV-cache offloading address real operational needs.
- Choose KubeRay/RayCluster or LWS when the serving design requires distributed inference across nodes, and validate hardware topology and autoscaling before rollout.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

