Reducing the Size of a Model-Serving Docker Image
10 min read · updated August 11, 2026
Most advice about shrinking Docker images was written for web applications, where switching to a slim base image is the whole answer. On a model-serving image the base is often a rounding error next to one RUN pip install. Find out which before you spend a day on the wrong layer.
Measure before you change anything
Two numbers, and they are different. The first is the uncompressed size on a host, which is what fills a node’s disk:
docker image ls --format '{{.Repository}}:{{.Tag}}\t{{.Size}}' | grep inferenceThe second is the compressed size that crosses the network on a pull, which is what shows up as cold-start latency and as registry egress cost. That is the sum of the compressed layer blobs, and you can read it from the registry rather than guessing at a compression ratio:
docker manifest inspect registry.example.com/inference:v7 \ | jq '[.layers[].size] | add / 1024 / 1024 | round'
Write both down. Every change below moves one of them more than the other, and a change that halves the on-disk size while leaving the compressed size alone has done nothing for your pull time. Weights and compiled libraries are already close to incompressible, so on an inference image the two numbers are usually closer together than on an ordinary application image.
Read the layers
docker history attributes bytes to the instruction that created them. This is the single most useful command in this whole exercise and it is the one people skip.
docker history --no-trunc --format '{{.Size}}\t{{.CreatedBy}}' inference:localRead it top to bottom and find the one or two lines that hold most of the total. On a GPU inference image the ranking is fairly predictable in shape, if not in magnitude:
- The Python dependency install. Modern PyTorch wheels for CUDA pull the CUDA runtime, cuDNN, cuBLAS and NCCL in as separate
nvidia-*packages, so onepip install torchinstalls a substantial fraction of a CUDA toolkit intosite-packages. - The CUDA base image, if you used one. The
develvariants carry the compiler and headers; theruntimevariants carry the shared libraries; thebasevariants carry neither. - Model weights, if you baked them. Usually the largest single item when present, and a decision with its own tradeoffs covered in baking versus mounting.
- Everything else, which is normally a small enough share that optimising it is not worth the review time.
The base image, which is rarely the problem
Two changes are worth making anyway because they are nearly free.
First, if you are on a full python:3.12, move to python:3.12-slim. The slim variants drop the build toolchain and much of the documentation and locale data. If your build needs a compiler, that belongs in a separate build stage rather than in the runtime base.
Second, if you are on an NVIDIA CUDA base, check which variant. NVIDIA publishes base, runtime and devel tags with increasing content, and images are frequently built on devel because that is what a Dockerfile found online used. If your service only runs pre-built kernels, the runtime tag is enough. If your Python dependencies already vendor their own CUDA libraries — which the PyTorch CUDA wheels do — you may not need a CUDA base image at all, only the driver injected by the NVIDIA Container Toolkit. That is the single largest reduction available on many inference images and it costs one line.
Alpine is the standard next suggestion and it is usually wrong here. Alpine uses musl rather than glibc, so the manylinux wheels that make Python data work installable do not apply, and pip falls back to building from source — slower builds, larger toolchain, and frequently a numeric library that will not build at all.
The dependency layer, which usually is
Three specific things, in order of how much they typically return.
- Install the right framework build. PyTorch publishes CPU-only wheels on a separate index. If the container will never touch a GPU — a tokeniser sidecar, an embedding service on CPU, a preprocessing worker — installing the default CUDA-enabled wheel ships an entire GPU stack that will never execute. Pin the index with
--index-urlfor the CPU build and confirm afterwards thattorch.version.cudaisNone. - Prune what you are actually importing. Requirements files accumulate. Generate the real import set from the source tree and compare it against the installed set rather than reasoning about it — a single unused dependency that pulls a plotting or notebook stack is often tens of megabytes of transitive install.
- Do not let pip keep its cache in the layer.
pip install --no-cache-dirstops the downloaded wheels being written under~/.cache/pipinside the image. Use this or a BuildKit cache mount, and understand the difference:--no-cache-dirmakes the image smaller and makes rebuilds re-download, while--mount=type=cachekeeps the downloads outside the image entirely and gives you both. The mount is strictly better when your builder is persistent, and does nothing when it is not.
Deletions that do not delete
Layers are additive and immutable. Removing a file in a later instruction records a whitelist entry saying the file is gone; the bytes stay in the earlier layer and are still pulled. This is why the classic apt-get idiom chains everything into one instruction:
RUN apt-get update \ && apt-get install -y --no-install-recommends libgomp1 \ && rm -rf /var/lib/apt/lists/*
Split across three RUN lines this saves nothing at all, because the package lists were already committed to a layer before the rm ran. The same trap catches a RUN that downloads an archive, extracts it and deletes the archive across separate instructions.
--no-install-recommends is worth its own sentence: apt installs recommended packages by default, and on a minimal base that routinely doubles what a single install pulls in. It is one flag and it is frequently the second-largest win on the page.
Re-measure and decide when to stop
Run the same two commands from the top and compare against the numbers you wrote down. Then stop at the point where the remaining layers are all small, because past that the returns fall off a cliff and every further change costs build complexity that somebody maintains forever.
It is worth being clear about what the size is buying. Below a certain point image size stops driving cold starts because the node has the layers cached already and the real startup cost is loading weights into GPU memory. Above it, size drives pull time on every scale-out event, registry storage, egress bills, and the failure modes in pushing a large image. Know which side of that line you are on before deciding how much more effort the next hundred megabytes deserves.