Skip to content

Reducing the Size of a Model-Serving Docker Image

10 min read · updated August 11, 2026

Most advice about shrinking Docker images was written for web applications, where switching to a slim base image is the whole answer. On a model-serving image the base is often a rounding error next to one RUN pip install. Find out which before you spend a day on the wrong layer.

Measure before you change anything

Two numbers, and they are different. The first is the uncompressed size on a host, which is what fills a node’s disk:

docker image ls --format '{{.Repository}}:{{.Tag}}\t{{.Size}}' | grep inference

The second is the compressed size that crosses the network on a pull, which is what shows up as cold-start latency and as registry egress cost. That is the sum of the compressed layer blobs, and you can read it from the registry rather than guessing at a compression ratio:

docker manifest inspect registry.example.com/inference:v7 \
  | jq '[.layers[].size] | add / 1024 / 1024 | round'

Write both down. Every change below moves one of them more than the other, and a change that halves the on-disk size while leaving the compressed size alone has done nothing for your pull time. Weights and compiled libraries are already close to incompressible, so on an inference image the two numbers are usually closer together than on an ordinary application image.

Read the layers

docker history attributes bytes to the instruction that created them. This is the single most useful command in this whole exercise and it is the one people skip.

docker history --no-trunc --format '{{.Size}}\t{{.CreatedBy}}' inference:local

Read it top to bottom and find the one or two lines that hold most of the total. On a GPU inference image the ranking is fairly predictable in shape, if not in magnitude:

  • The Python dependency install. Modern PyTorch wheels for CUDA pull the CUDA runtime, cuDNN, cuBLAS and NCCL in as separate nvidia-* packages, so one pip install torch installs a substantial fraction of a CUDA toolkit into site-packages.
  • The CUDA base image, if you used one. The devel variants carry the compiler and headers; the runtime variants carry the shared libraries; the base variants carry neither.
  • Model weights, if you baked them. Usually the largest single item when present, and a decision with its own tradeoffs covered in baking versus mounting.
  • Everything else, which is normally a small enough share that optimising it is not worth the review time.
Absolute sizes for base images and framework wheels change with every CUDA and framework release, which is why none are quoted here. Run the command against your own image; the number you get is the only one that is true for your build.

The base image, which is rarely the problem

Two changes are worth making anyway because they are nearly free.

First, if you are on a full python:3.12, move to python:3.12-slim. The slim variants drop the build toolchain and much of the documentation and locale data. If your build needs a compiler, that belongs in a separate build stage rather than in the runtime base.

Second, if you are on an NVIDIA CUDA base, check which variant. NVIDIA publishes base, runtime and devel tags with increasing content, and images are frequently built on devel because that is what a Dockerfile found online used. If your service only runs pre-built kernels, the runtime tag is enough. If your Python dependencies already vendor their own CUDA libraries — which the PyTorch CUDA wheels do — you may not need a CUDA base image at all, only the driver injected by the NVIDIA Container Toolkit. That is the single largest reduction available on many inference images and it costs one line.

Alpine is the standard next suggestion and it is usually wrong here. Alpine uses musl rather than glibc, so the manylinux wheels that make Python data work installable do not apply, and pip falls back to building from source — slower builds, larger toolchain, and frequently a numeric library that will not build at all.

The dependency layer, which usually is

Three specific things, in order of how much they typically return.

  1. Install the right framework build. PyTorch publishes CPU-only wheels on a separate index. If the container will never touch a GPU — a tokeniser sidecar, an embedding service on CPU, a preprocessing worker — installing the default CUDA-enabled wheel ships an entire GPU stack that will never execute. Pin the index with --index-url for the CPU build and confirm afterwards that torch.version.cuda is None.
  2. Prune what you are actually importing. Requirements files accumulate. Generate the real import set from the source tree and compare it against the installed set rather than reasoning about it — a single unused dependency that pulls a plotting or notebook stack is often tens of megabytes of transitive install.
  3. Do not let pip keep its cache in the layer. pip install --no-cache-dir stops the downloaded wheels being written under ~/.cache/pip inside the image. Use this or a BuildKit cache mount, and understand the difference: --no-cache-dir makes the image smaller and makes rebuilds re-download, while --mount=type=cache keeps the downloads outside the image entirely and gives you both. The mount is strictly better when your builder is persistent, and does nothing when it is not.

Deletions that do not delete

Layers are additive and immutable. Removing a file in a later instruction records a whitelist entry saying the file is gone; the bytes stay in the earlier layer and are still pulled. This is why the classic apt-get idiom chains everything into one instruction:

RUN apt-get update \
 && apt-get install -y --no-install-recommends libgomp1 \
 && rm -rf /var/lib/apt/lists/*

Split across three RUN lines this saves nothing at all, because the package lists were already committed to a layer before the rm ran. The same trap catches a RUN that downloads an archive, extracts it and deletes the archive across separate instructions.

--no-install-recommends is worth its own sentence: apt installs recommended packages by default, and on a minimal base that routinely doubles what a single install pulls in. It is one flag and it is frequently the second-largest win on the page.

Re-measure and decide when to stop

Run the same two commands from the top and compare against the numbers you wrote down. Then stop at the point where the remaining layers are all small, because past that the returns fall off a cliff and every further change costs build complexity that somebody maintains forever.

It is worth being clear about what the size is buying. Below a certain point image size stops driving cold starts because the node has the layers cached already and the real startup cost is loading weights into GPU memory. Above it, size drives pull time on every scale-out event, registry storage, egress bills, and the failure modes in pushing a large image. Know which side of that line you are on before deciding how much more effort the next hundred megabytes deserves.