Fixing CrashLoopBackOff on a GPU Inference Pod
9 min read · updated August 11, 2026
CrashLoopBackOff is not a diagnosis. It says the container started, exited, and the kubelet is waiting before starting it again. On a GPU pod there are four common reasons underneath it, and the container’s exit code tells you which one you have in about ten seconds — long before anyone needs to guess at YAML.
What the status actually says
The Kubernetes documentation is explicit that CrashLoopBackOff is a display value in kubectl’s STATUS column, not a pod phase. It means the container has terminated and the kubelet is in its restart backoff. The interesting fact is not that it is restarting; it is why the previous instance stopped, which is recorded on the container status as lastState.terminated and is overwritten each time it happens.
A GPU inference pod fails at four distinct points, and they look identical from the pod list: the process ran out of memory, the CUDA libraries in the image did not match the driver on the node, the container never got a GPU at all, or the model server exited on its own argument validation. Each of these has a different signature in the termination record.
Get the exit code before you change anything
Two commands. The first reads the structured termination record for the previous instance; the second reads what the process printed before it died, which kubectl logs without --previous will not show you because the current instance may still be starting.
kubectl get pod inference-0 -n serving \
-o jsonpath='{.status.containerStatuses[0].lastState.terminated}' | jq
kubectl logs inference-0 -n serving --previous --tail=80The first command returns an object with exitCode, reason, signal, startedAt and finishedAt. Read exitCode and the gap between the two timestamps together: an exit at four seconds and an exit at four minutes are different bugs even when the code is the same.
Exit codes above 128 are signals. The convention is 128 plus the signal number, so 137 is SIGKILL (9), 143 is SIGTERM (15) and 139 is SIGSEGV (11). Anything below 128 came from the process itself: 1 is a generic application error, 127 is “command not found” from the shell, and 126 is “found but not executable”.
137: something killed it for memory
Exit code 137 with reason: OOMKilled means the kernel’s OOM killer terminated the process because the container exceeded its cgroup memory limit. Kubernetes sets that reason field; the application never sees it, which is why the logs often end mid-sentence with nothing that looks like an error.
On an inference pod this is nearly always host RAM rather than GPU memory, and it is nearly always the weight load. Loading a checkpoint typically materialises tensors in CPU memory before moving them to the device, so the peak host requirement during startup can approach the size of the checkpoint on disk even though steady-state usage afterwards is small. A pod whose resources.limits.memory was sized from a running container will be killed every time it is restarted cold.
The distinguishing test is timing: an OOM during weight load happens at a repeatable point in the first minute, and the log’s last line is always the same. Raise limits.memory above the peak, not above the steady state. If the checkpoint is memory-mapped from a volume rather than read into a buffer, the peak drops sharply — that is part of why where the weights live is a runtime decision and not only a packaging one.
torch.OutOfMemoryError: CUDA out of memory — because the allocation failed inside the process rather than the process being killed from outside it.1 or 127: the CUDA stack did not load
Exit code 1 with a traceback is the image and the node disagreeing. The string to look for in --previous logs is CUDA driver version is insufficient for CUDA runtime version. That is cudaErrorInsufficientDriver, and it means the CUDA runtime compiled into your image is newer than the driver installed on the node. The direction matters: NVIDIA’s drivers are backward-compatible with older runtimes, so an old image on a new node is usually fine and a new image on an old node is not.
The neighbouring failure is a linker error rather than a CUDA error: libcuda.so.1: cannot open shared object file, or 127 from an entrypoint that shells out to nvidia-smi and does not find it. Both mean the driver libraries were not injected into the container. That is not a CUDA version problem — it is the next section.
Confirm the mismatch by asking the node rather than the pod. Schedule a throwaway pod with the same node selector and run nvidia-smi, whose header prints the installed driver version and the maximum CUDA version that driver supports. Compare that to the CUDA version in your base image tag. If a cluster has mixed node images — which is normal after a partial node-pool upgrade — the pod will crash on some nodes and run on others, and that intermittency is the tell. Pin the workload with node affinity on the GPU label rather than hoping the scheduler picks a good one.
The GPU was never attached
If the container starts on a GPU node but sees no device, check that the pod actually requested one. A GPU is an extended resource, and a container that does not name it in resources.limits is not given one:
resources:
limits:
nvidia.com/gpu: 1Extended resources must appear under limits; Kubernetes copies the value to requests automatically, and the two may not differ. A pod that omits this schedules happily onto a GPU node and then fails in the container, which is exactly the confusing case — the pod is not stuck Pending, so nothing points at scheduling. If the pod is stuck Pending instead, the problem is capacity or taints and the page you want is insufficient nvidia.com/gpu.
The other half of this is cluster-side: the NVIDIA device plugin DaemonSet is what advertises nvidia.com/gpu and what mounts the driver libraries into the container. Run kubectl get nodes -o jsonpath over status.allocatable and confirm the resource exists on the node at all. If it does not, no pod on that node will ever see a device, however the manifest is written.
Why the restarts get slower
The kubelet restarts a failed container with exponential backoff. The Kubernetes documentation describes an initial delay that doubles on each successive failure up to a maximum of five minutes, and the backoff resets once a container has run successfully for ten minutes. The backoff is tracked per container, not per pod.
Two practical consequences. First, a pod that has been failing for an hour is retrying every five minutes, so after you push a fix you may wait minutes for the next attempt — delete the pod rather than concluding the fix did not work. Second, the ten-minute reset window is why a model server that takes eight minutes to load weights and then crashes never escapes the backoff: it never accumulates enough successful runtime to reset. Fix the crash; do not tune the backoff, which is not a knob on the pod spec.
Once it stays up, the remaining trap is that a running container is not a ready one. A model server accepts TCP connections long before the weights are resident, so traffic arrives to a process that cannot serve it. That is what a readiness probe tied to model load exists for, and it is a separate fix from this one.