Deploying Inference on AWS Fargate
9 min read · updated August 11, 2026
Fargate removes the container instance, which removes the GPU with it. That is a real constraint and it is also narrower than it sounds: a large share of what people call an inference service is a CPU-bound process that calls a model somewhere else.
There is no GPU on Fargate, and why
AWS lists gpu among the task definition parameters that are not valid in Fargate tasks, alongside privileged, ipcMode, devices, placementConstraints and dockerVolumeConfiguration. That list is the shape of the whole service rather than a gap in it. Fargate’s bargain is that you never see the host: no AMI, no agent configuration, no driver version, no instance type. A GPU cannot be handed to a container without a host-side driver stack whose version has to match the CUDA build inside your image, and that coupling is precisely the thing Fargate exists to hide. Which is why the answer is not “not yet” so much as “not without giving back the property you chose it for”.
The practical consequence: if your container holds model weights and runs them, you want the EC2 launch type — see running a GPU task on Amazon ECS. If your container tokenises, retrieves, calls an API, re-ranks with a small cross-encoder and assembles a response, Fargate is a good fit and the rest of this page is about sizing it.
The sizing grid you are choosing from
Fargate does not let you pick arbitrary CPU and memory. AWS documents a fixed grid of task-level combinations, and a request outside it is rejected at RegisterTaskDefinition rather than at run time. At the time of writing the documented combinations are 256 CPU units (.25 vCPU) with 512 MiB, 1 GB or 2 GB; 512 (.5 vCPU) with 1–4 GB; 1024 (1 vCPU) with 2–8 GB; 2048 (2 vCPU) with 4–16 GB in 1 GB steps; 4096 (4 vCPU) with 8–30 GB in 1 GB steps; 8192 (8 vCPU) with 16–60 GB in 4 GB steps; 16384 (16 vCPU) with 32–120 GB in 8 GB steps; and 32768 (32 vCPU) with 60 GB, 120 GB or 244 GB. The 8 vCPU tier and above require Linux platform version 1.4.0 or later.
For a CPU inference container the memory column is usually what binds, because a model loaded for CPU execution sits resident for the life of the task, and the grid forces you upward in whole vCPUs to reach it. Wanting 12 GB with half a core is not an option; 12 GB starts at 2 vCPU. That coupling is worth doing the arithmetic on before you write the definition, because it changes the price of the service more than any tuning you do afterwards.
A task definition that runs
Fargate always uses the awsvpc network mode, so the task gets its own elastic network interface and its own security group, and you name subnets at run time rather than in the definition:
{
"family": "rerank-api",
"requiresCompatibilities": ["FARGATE"],
"networkMode": "awsvpc",
"cpu": "2048",
"memory": "8192",
"runtimePlatform": {
"cpuArchitecture": "ARM64",
"operatingSystemFamily": "LINUX"
},
"executionRoleArn": "arn:aws:iam::111122223333:role/ecsTaskExecutionRole",
"taskRoleArn": "arn:aws:iam::111122223333:role/rerankTaskRole",
"containerDefinitions": [
{
"name": "api",
"image": "111122223333.dkr.ecr.us-east-1.amazonaws.com/rerank:2.1.0",
"essential": true,
"portMappings": [{ "containerPort": 8080, "protocol": "tcp" }],
"ulimits": [
{ "name": "nofile", "softLimit": 65535, "hardLimit": 65535 }
],
"logConfiguration": {
"logDriver": "awslogs",
"options": {
"awslogs-group": "/ecs/rerank-api",
"awslogs-region": "us-east-1",
"awslogs-stream-prefix": "ecs"
}
}
}
]
}- Register it with
aws ecs register-task-definition --cli-input-json file://taskdef.json. Passing--requires-compatibilities FARGATEis what makes ECS validate the definition against the Fargate rules at registration time rather than letting an invalid parameter surface as a failed task later. - Create the service with an explicit network configuration:
aws ecs create-service --cluster app --service-name rerank-api --task-definition rerank-api --desired-count 2 --launch-type FARGATE --network-configuration "awsvpcConfiguration={subnets=[subnet-aaa,subnet-bbb],securityGroups=[sg-111],assignPublicIp=DISABLED}". - Give the task somewhere to pull from. With
assignPublicIp=DISABLEDin a private subnet, the image pull needs either a NAT gateway or ECR interface VPC endpoints. This is the single most common first-deploy failure on Fargate, and it presents as a task that stops with a CannotPullContainerError before your code has run once. - Attach a load balancer target group of type
ip, notinstance. Underawsvpcthe target is the task’s ENI address.
The ulimits block is worth keeping. AWS documents the Fargate default nofile soft and hard limits as 65535, raisable to 1048576 — and a Python service holding many concurrent upstream HTTP connections is exactly the shape that hits a file descriptor ceiling and reports it as a connection error.
Image pull is most of your start time
On Fargate there is no warm local image cache to land on: each task starts on fresh capacity and pulls. For a container carrying a PyTorch runtime and a few gigabytes of weights, the pull dominates everything else about task start, and it is charged as time your service is not yet serving.
AWS documents Seekable OCI as the mitigation. On Linux platform version 1.4.0, if a SOCI index exists alongside the image in a private ECR repository, Fargate starts the container while the rest of the image downloads in the background, rather than waiting for the whole image. AWS notes it is worth trying on images larger than roughly 250 MiB compressed and that zstd-compressed images are not supported. Fargate also publishes per-phase timestamps on the task, so you can read image-pull duration directly out of describe-tasks instead of inferring it — do that before optimising anything, because the answer is often that a base image is carrying build tooling into production.
What Fargate is genuinely good at here
- The stateless request layer in front of a model. Auth, prompt assembly, retrieval, response validation, logging. All CPU and network, none of it wanting a GPU.
- Queue workers that spend their life waiting. A worker whose latency is dominated by an upstream model call is idle CPU with an open socket. Paying for a GPU to hold that socket is the expensive mistake; scaling it on queue depth rather than CPU is the fix.
- Small CPU models. Embedding models in the tens of millions of parameters, classifiers and cross-encoder re-rankers run acceptably on 2–4 vCPU, and the ARM64 option on Fargate is usually the cheaper one for them.
- Anything with spiky, low-duty-cycle traffic. No instance to keep warm and no capacity provider to reason about. The trade is the pull time above, which is why the image size matters more here than on EC2.