Skip to content

AWS Batch for Long-Running Model Jobs

10 min read · updated August 11, 2026

AWS documents Lambda’s maximum function timeout as 900 seconds. Embedding a document corpus, running an evaluation suite, or summarising a month of transcripts does not fit in fifteen minutes, and the usual workarounds — chunking into self-invoking Lambdas, a Step Functions loop — are ways of avoiding a service that already exists for exactly this.

When Lambda runs out

The 900-second ceiling is the obvious constraint, but it is rarely the only one that has been reached by the time someone goes looking for Batch. The others: 10,240 MB is the memory maximum, the ephemeral /tmp directory tops out at 10,240 MB, and a synchronous invocation payload is capped at 6 MB in each direction. A job that downloads a corpus, holds it in memory and writes a large artifact bumps into all four.

Batch removes all of them by running your container on ECS or EKS capacity it manages: no duration limit, whatever instance size you are willing to pay for, and the job reads and writes S3 rather than passing payloads. What you give up is the invocation model — a Batch job is submitted and polled, not called and awaited.

Three resources, in this order

Batch is three objects, and they must be created in dependency order because each references the last.

  • Compute environment — the capacity. Its type is MANAGED or UNMANAGED; within computeResources the type is EC2, SPOT, FARGATE or FARGATE_SPOT. Managed plus Fargate is the lowest-operations option and needs no instance role or AMI decisions.
  • Job queue — where jobs wait. It carries a priority integer where higher is evaluated first, a state of ENABLED or DISABLED, and a computeEnvironmentOrder array whose order is the placement preference. AWS documents one constraint that catches people: all compute environments attached to a queue must be either EC2-family (EC2 or SPOT) or Fargate-family (FARGATE or FARGATE_SPOT) — the two cannot be mixed.
  • Job definition — the container and its resources. Versioned: registering the same name again produces revision 2, and a submitted job can name a specific revision.
aws batch create-compute-environment \
  --compute-environment-name embed-fargate \
  --type MANAGED \
  --state ENABLED \
  --compute-resources '{
    "type": "FARGATE",
    "maxvCpus": 64,
    "subnets": ["subnet-0aaa111","subnet-0bbb222"],
    "securityGroupIds": ["sg-0ccc333"]
  }'

aws batch create-job-queue \
  --job-queue-name embed-queue \
  --state ENABLED \
  --priority 10 \
  --compute-environment-order order=1,computeEnvironment=embed-fargate

maxvCpus is the ceiling on concurrent capacity and therefore your real spend control — a queue with a hundred jobs and a maxvCpus of 64 runs them in waves rather than all at once. Note also that AWS documents a default Fargate On-Demand vCPU quota of 6 for the account, which is a separate limit from maxvCpus and the one that will actually stop you first. Check it in the Service Quotas console before assuming a compute environment is misconfigured.

Registering the job definition

  1. Push your container to ECR. It needs no Batch-specific code — unlike a Lambda container image, which must implement the runtime interface, a Batch container is an ordinary image whose process exits 0 on success.
  2. Decide the two roles. executionRoleArn is what pulls the image and writes logs, and is required for Fargate jobs. jobRoleArn is what your code uses — the one that needs bedrock:InvokeModel and S3 access. Conflating them is the most common Batch IAM error.
  3. Express CPU and memory through resourceRequirements. The vcpus and memory fields still parse and are documented as deprecated.
{
  "jobDefinitionName": "embed-corpus",
  "type": "container",
  "platformCapabilities": ["FARGATE"],
  "containerProperties": {
    "image": "111122223333.dkr.ecr.us-east-1.amazonaws.com/embed:1.4.0",
    "command": ["python", "embed.py", "--manifest", "Ref::manifest"],
    "jobRoleArn": "arn:aws:iam::111122223333:role/embed-job-role",
    "executionRoleArn": "arn:aws:iam::111122223333:role/batch-execution-role",
    "resourceRequirements": [
      { "type": "VCPU",   "value": "4" },
      { "type": "MEMORY", "value": "8192" }
    ]
  },
  "timeout": { "attemptDurationSeconds": 10800 },
  "retryStrategy": {
    "attempts": 3,
    "evaluateOnExit": [
      { "onStatusReason": "Host EC2*", "action": "RETRY" },
      { "onExitCode": "1", "action": "EXIT" }
    ]
  }
}

On Fargate the VCPU value must be one of the supported values — AWS lists 0.25, 0.5, 1, 2, 4, 8, 16 and 32 — and the MEMORY value must be one that is supported for that vCPU count. A memory figure that is legal for 4 vCPU and not for 2 fails at registration with a message about supported values, which is confusing until you know the two are coupled. For GPU work, add a { "type": "GPU", "value": "1" } requirement and use an EC2 compute environment; Fargate has no GPUs.

Ref::manifest in the command is Batch’s parameter substitution: the value comes from parameters at submission time, so one job definition serves many inputs without rebuilding the image.

Submitting and watching

aws batch submit-job \
  --job-name embed-2026-08-11 \
  --job-queue embed-queue \
  --job-definition embed-corpus \
  --parameters manifest=s3://my-bucket/manifests/2026-08-11.json \
  --container-overrides '{"environment":[{"name":"BATCH_SIZE","value":"64"}]}'

The response carries a jobId. A job then moves through SUBMITTED, PENDING, RUNNABLE, STARTING, RUNNING and finally SUCCEEDED or FAILED. Learn one of these: RUNNABLE means Batch has accepted the job and cannot place it. A job stuck there is not slow, it is blocked — almost always by maxvCpus, an account quota, a resource request larger than any instance in the environment, or a subnet with no route to ECR. aws batch describe-jobs --jobs ID returns a statusReason that usually names which.

There is also a guard for this: jobStateTimeLimitActions on the job queue lets you tell Batch to act on jobs that sit at the head of the queue in a given state beyond maxTimeSeconds, so a job wedged in RUNNABLE fails loudly rather than waiting forever.

Timeouts, retries and array jobs

Long jobs fail differently from short ones: the cost of a failure at minute 170 of a three-hour run is the whole run. Three fields decide what that costs you.

  • timeout.attemptDurationSeconds — measured from the attempt’s startedAt, minimum 60 seconds. Always set it. Without one, a container that hangs on a socket holds capacity until somebody notices. For array jobs the timeout applies to the children, not the parent.
  • retryStrategy.attempts — between 1 and 10. Blind retries on a job that consumes model tokens re-spend them, which is why evaluateOnExit matters more here than in most workloads: up to five conditions matching on onExitCode, onReason or onStatusReason, each with a RETRY or EXIT action. Retry an interrupted Spot instance; do not retry your own exit code 1. Note that if evaluateOnExit is specified and nothing matches, the job is retried — the default is permissive, so make the last entry an explicit EXIT.
  • Array jobs — --array-properties size=500 submits one parent and 500 children, each with AWS_BATCH_JOB_ARRAY_INDEX in its environment. This is the right shape for embedding a corpus: a failed shard retries alone instead of restarting the corpus, and the index maps to a slice of the manifest.

Make the container idempotent per shard — check whether the output object already exists in S3 before doing the work. Combined with array jobs it turns a retry from an expensive redo into a cheap no-op, which is what makes a three-attempt strategy affordable at all. If your long job is specifically offline inference against a hosted model, check Bedrock batch inference first: it is a different service with its own discount, and running it yourself on Batch is worth doing only when you need the container.