Fixing a Push Timeout on a Large AI Docker Image
10 min read · updated August 11, 2026
A large image push fails in four distinguishable ways, and each points at a different component. Read the exact string before changing anything, because “retry with a better connection” fixes exactly one of them.
Which error did you get
413 Request Entity Too Large— something in the HTTP path refused the body by size. This is a proxy or ingress limit almost every time, not the registry itself.net/http: TLS handshake timeoutor ani/o timeouton a read — the connection stalled. Genuine network path problem, or a proxy read timeout shorter than the time a large blob takes to upload.blob upload unknownorblob upload invalid— the registry lost the in-progress upload session. Retrying restarts that blob from zero, which on a huge layer can loop forever.unauthorized: authentication requiredpartway through a push that started fine — the token expired during the upload. See below; this one is unrelated to size except that size is what made the push take long enough.
The push is not one request. The client uploads each layer blob separately, then a final manifest. So a push that gets 80% through and fails has completed some blobs successfully — which is useful, because those are already in the registry and a retry skips them.
Check the layer sizes first
Almost everything below turns on the size of the largest single layer rather than the total image size. A 40 GB image made of forty 1 GB layers behaves completely differently from a 40 GB image with one 38 GB layer.
docker history --no-trunc --format '{{.Size}}\t{{.CreatedBy}}' myimage:v7 | sort -h | tail -5If one layer holds most of the image, that is your problem statement. On an inference image it is usually either a baked checkpoint or a single pip install of the framework. Splitting it — copying weights in several COPY instructions, or separating the heavy framework install from the rest of the requirements — is the highest-leverage change available and it also makes retries cheap, because a failure now costs one small blob rather than the whole thing.
Registry limits are real and published
Registries publish hard limits and they are not always where you expect. Amazon documents ECR’s in its service quotas page: a Maximum layer size of 52,000 MiB, a Layer parts quota of 4,200 and a Maximum layer part size of 10 MiB, none of which are adjustable at the time of writing. The part quotas apply when you drive the multipart upload through the ECR API directly rather than through docker push.
The rate quotas on the same page are the ones that catch parallel CI. Amazon documents Rate of InitiateLayerUpload requests at 100 per second and Rate of CompleteLayerUpload requests at 100 per second, both adjustable via Service Quotas, alongside Rate of UploadLayerPart requests at 500 per second. Twenty concurrent build jobs each pushing a many-layered image can reach these, and the failure looks like a throttle rather than a size problem.
Amazon ECR service quotasThe proxy in front of the registry
A 413 that only happens on a self-hosted or on-premises registry is nearly always a reverse proxy. nginx defaults client_max_body_size to a value far below any container layer, so the registry never sees the request. The fix is on the proxy, not the registry: set client_max_body_size 0 on the registry location to disable the check, and raise proxy_read_timeout and proxy_send_timeout at the same time, because the size limit and the read timeout usually need fixing together and fixing one alone just changes which error you get.
The equivalent for a Kubernetes ingress is an annotation on the Ingress object, and for other proxies a request-body-size directive under a different name. Whatever the proxy, the diagnostic is the same: push directly to the registry’s own address, bypassing the proxy. If that works, the proxy is the problem and no amount of registry configuration will help.
One more path worth ruling out: some CDNs and edge proxies impose their own request body limit, which is why a push can start failing after somebody puts the registry hostname behind one for TLS. Nothing in the registry changed.
Credentials expiring mid-push
Registry credentials are frequently short-lived tokens. An ECR login token obtained from GetAuthorizationToken is valid for a bounded period, and a push that takes longer than the remaining validity fails part-way with unauthorized: authentication required even though the same command worked yesterday on a smaller image.
Two things follow. Log in immediately before the push rather than at the start of a long pipeline — aws ecr get-login-password | docker login --password-stdin costs nothing and removes the whole class of failure. And treat a mid-push authorisation error as a clock problem rather than a permissions problem: if the credential were wrong, the first blob would have failed, not the fifteenth.
Make the push smaller
- Split the largest layer. Several medium layers upload in parallel and retry cheaply; one enormous layer does neither.
- Push from inside the same network as the registry. A CI runner in the same region as the registry removes both the latency and most of the timeout risk. Pushing 30 GB from a laptop over a consumer uplink will fail eventually whatever you configure.
- Stop shipping the bytes at all. If the large layer is model weights, the question is whether they belong in the image at all — baking versus mounting is the decision, and mounting removes this failure mode entirely rather than mitigating it.
- Reduce what is there. Everything in reducing a model-serving image applies, and a change that removes a CUDA toolkit you were not using is worth more than any transport tuning.
- Push once, replicate afterwards. If the same image must exist in several regions, use the registry’s own replication rather than pushing from CI to each one. Registry-side copies do not traverse your build network and do not fail on your timeouts.