Deploying vLLM on Kubernetes
10 min read · updated August 11, 2026
vLLM ships an official image whose entrypoint is an OpenAI-compatible HTTP server. Getting it onto Kubernetes is a short manifest; getting it to stay up under a rollout is three fields that the short manifest leaves out.
What the container actually is
The image published as vllm/vllm-openai starts an HTTP server that implements the OpenAI API surface — /v1/completions, /v1/chat/completions, /v1/models — on port 8000, plus a /health endpoint for probes. The vLLM project’s own Kubernetes deployment guide uses exactly this image and port. The command is vllm serve MODEL_ID followed by flags.
Two things follow from that shape. First, the model is a runtime argument, not part of the image, so the container downloads weights on start unless you have mounted them; a seven-billion-parameter model in bf16 is roughly fourteen gigabytes of download before the server binds its port. Second, because the API surface is OpenAI-shaped, any client library you already use points at it by changing a base URL — which is the entire reason to run vLLM rather than a bespoke server.
The download is worth planning for rather than tolerating. An emptyDir cache, as in the manifest below, means the weights survive a container restart but not a pod reschedule, so a node failure costs the full download again. A PersistentVolumeClaim on a shared filesystem shares one copy across replicas and turns the download into a one-off; baking the weights into the image trades registry storage and image pull time for a deterministic start. Which of the three is right depends on how often pods move, and the answer is almost never “whatever the example manifest did”.
The Deployment
apiVersion: apps/v1
kind: Deployment
metadata:
name: mistral-7b
namespace: inference
spec:
replicas: 1
selector:
matchLabels:
app: mistral-7b
template:
metadata:
labels:
app: mistral-7b
spec:
terminationGracePeriodSeconds: 120
volumes:
- name: shm
emptyDir:
medium: Memory
sizeLimit: 2Gi
- name: hf-cache
emptyDir: {}
containers:
- name: vllm
image: vllm/vllm-openai:<pin a tag>
command: ["/bin/sh", "-c"]
args:
- "vllm serve mistralai/Mistral-7B-Instruct-v0.3 --port 8000"
ports:
- containerPort: 8000
env:
- name: HF_TOKEN
valueFrom:
secretKeyRef:
name: hf-token-secret
key: token
resources:
limits:
cpu: "8"
memory: 32Gi
nvidia.com/gpu: "1"
startupProbe:
httpGet:
path: /health
port: 8000
periodSeconds: 10
failureThreshold: 60
readinessProbe:
httpGet:
path: /health
port: 8000
periodSeconds: 5
volumeMounts:
- name: shm
mountPath: /dev/shm
- name: hf-cache
mountPath: /root/.cache/huggingfacelatest is not a tag you want here. vLLM releases frequently and flag names have changed between releases; pin a version and upgrade deliberately.The three fields everyone omits
The shared memory volume. A container gets a small /dev/shm by default — 64 MiB under the usual runtime defaults. vLLM and PyTorch use shared memory for inter-process communication, and the failure when it runs out is a torch error about a bus or a shared-memory allocation, several layers below anything that says “increase shm”. The fix is the emptyDir with medium: Memory mounted at /dev/shm, which the vLLM Kubernetes guide includes for this reason. Note that a memory-backed emptyDir counts against the pod’s memory limit, so budget for it.
The startup probe. A readiness probe with a long initialDelaySeconds is the common workaround for slow model loading and it is the wrong one, because the same delay then applies to every restart and you cannot make it long enough for a cold cache without making failure detection useless afterwards. A startup probe exists precisely for this: it suppresses the other probes until it passes once, then hands over. failureThreshold: 60 at a ten-second period gives ten minutes to load, after which the readiness probe takes over at five seconds.
The grace period. The default terminationGracePeriodSeconds is 30. A generation request that is streaming 800 tokens can outlive that, so a rolling update truncates responses that were mid-flight. Raising it is necessary but not sufficient — see draining a GPU node without killing in-flight requests for the preStop hook that makes the grace period actually get used.
Service and a verified request
apiVersion: v1
kind: Service
metadata:
name: mistral-7b
namespace: inference
spec:
type: ClusterIP
selector:
app: mistral-7b
ports:
- protocol: TCP
port: 80
targetPort: 8000- Wait for readiness:
kubectl rollout status deploy/mistral-7b -n inference. Until the startup probe passes this will sit and that is correct behaviour, not a hang. - Confirm the model actually loaded, not just that the port opened:
kubectl exec deploy/mistral-7b -n inference -- curl -s localhost:8000/v1/models. The response lists the model id the server is serving. - Send a real completion from inside the cluster. The vLLM documentation’s own verification uses the completions endpoint against the Service DNS name:
curl http://mistral-7b.inference.svc.cluster.local/v1/completions \ -H "Content-Type: application/json" \ -d '{"model": "mistralai/Mistral-7B-Instruct-v0.3", "prompt": "San Francisco is a", "max_tokens": 7, "temperature": 0}' - A 404 with a message about the model not being found means the
modelfield in the body does not match the id the server registered. It must be the full repository id you passed tovllm serve, not the Kubernetes object name.
The server flags that change behaviour
--gpu-memory-utilization— the fraction of GPU memory vLLM claims for weights plus its KV cache. It preallocates, so a value near the top of the range makes a second process on the same GPU fail immediately rather than gradually. Lower it when you are sharing a device via time-slicing.--max-model-len— caps context length. Left unset, vLLM uses the model’s full advertised context, and on a smaller card the KV cache for that length may not fit; the server exits at startup with an explicit message about how much memory it needed. Setting this is usually the fix, not buying a bigger GPU.--tensor-parallel-size— shards one model across several GPUs in the same pod. It must match thenvidia.com/gpulimit exactly; a mismatch produces a startup failure that reads as a device-count error rather than a manifest error.--served-model-name— decouples the name clients send in themodelfield from the repository id the server loaded. Worth setting from the start: without it, every client hard-codes a Hugging Face path, and swapping the underlying model becomes a change to every caller instead of a change to one Deployment argument.
When it restarts and it is not the GPU
Three restart causes account for most of the confusion here, and only one of them involves the accelerator at all.
Exit code 137 with reason OOMKilled. That is the kernel killing the container for exceeding its host memory limit, not a CUDA out-of-memory error. It happens during weight loading, because the loader materialises tensors in host RAM before moving them to the device, and because the memory-backed /dev/shm volume counts against the same limit. A model whose checkpoint is 14 GB with a 16 GB memory limit and a 2 GB shm volume has no headroom at all. Check with kubectl describe pod POD and read Last State: Terminated, Reason: OOMKilled; a CUDA OOM instead appears in the container logs as a Python traceback and exits with 1.
Restarts before the server ever listens. If the liveness probe is running during model load, the kubelet kills a container that was going to work, and the pod loops forever getting slightly further each time. This is why the manifest above uses a startup probe: it is the only probe that suppresses the others.
ImagePullBackOff on a fresh node. The vLLM image is large, and on a node that has never pulled it the pull can exceed the kubelet’s image pull progress deadline under a slow registry. This is a node-level problem rather than a pod-level one, and it shows up predictably every time a spot node is replaced.