Skip to content

Running a Batch Inference Job on Bedrock

10 min read · updated August 11, 2026

Batch inference on Bedrock is a file in, a file out, and an asynchronous job between them. The interesting part is not submitting it — that is four parameters — but the output contract: records come back out of order, failures are per record rather than per job, and the summary you actually want is in a file most tutorials never mention.

When batch is the right shape

Batch trades latency for price. AWS’s Bedrock pricing page describes batch as offering a 50% lower price than on-demand inference for supported foundation models, which makes anything that does not need an answer this second — nightly enrichment, backfilling classifications, evaluating a prompt change over ten thousand historical inputs — a candidate.

Two documented limitations decide it more often than price does. AWS states that batch inference does not support tool calling or structured output via response_format, because each record is processed independently with no back-and-forth. And batch is not supported for provisioned models. So a pipeline built around a tool loop cannot simply be moved to batch; it has to be flattened into single-shot prompts first.

The input format

Input is one or more .jsonl files in S3. Each line is a JSON object with a recordId and a modelInput. The shape of modelInput depends on the invocation type you choose when you create the job: for the default InvokeModel type it must match that model’s body, and for the Converse type it must match a Converse request body.

{"recordId":"CALL0000001","modelInput":{"messages":[{"role":"user","content":[{"text":"Summarise this call transcript: ..."}]}],"inferenceConfig":{"maxTokens":1024}}}
{"recordId":"CALL0000002","modelInput":{"messages":[{"role":"user","content":[{"text":"Summarise this call transcript: ..."}]}],"inferenceConfig":{"maxTokens":1024}}}

Use the Converse type unless you have a reason not to. It means the record body is the same JSON your synchronous path already builds, so a batch job and a live call can share one prompt-construction function instead of drifting apart.

recordId is technically optional — AWS notes that if you omit it, Bedrock adds one in the output — but omitting it throws away your only join key back to your own data, and the output order will not save you. Make it your primary key.

If your prompts reference S3 objects (a video or image by s3Location rather than by bytes), AWS requires all of those resources to live in the same bucket and folder, and inputDataConfig must point at the folder containing them rather than at the individual .jsonl. S3 paths are case-sensitive here, which produces a job that starts and then fails on every record.

Five quotas govern the input and all five are worth reading rather than guessing: Minimum number of records per batch inference job, Records per input file per batch inference job, Records per batch inference job, Batch inference input file size and Batch inference job size. Those are the exact quota names to search for in the Amazon Bedrock endpoints and quotas reference. There is a minimum as well as a maximum, which is the one that catches people testing with three records.

Submitting the job

  1. Upload your .jsonl files to an input prefix, and create an output prefix. They should be different prefixes; the job writes several files and you do not want them interleaved with next night’s input.
  2. Create a service role Bedrock can assume, with s3:GetObject and s3:ListBucket on the input location and s3:PutObject on the output location. The trust policy is for bedrock.amazonaws.com; scope it with aws:SourceAccount and an aws:SourceArn condition on the job ARN pattern so the role cannot be used from elsewhere.
  3. Submit with CreateModelInvocationJob, naming the role, the model and both S3 locations.
  4. Record the returned jobArn. It is how you get status, and it is also what you tag for cost allocation if you run batch for several teams.
aws bedrock create-model-invocation-job \
  --job-name nightly-transcript-summaries-2026-08-11 \
  --role-arn arn:aws:iam::111122223333:role/BedrockBatchInferenceRole \
  --model-id us.anthropic.claude-sonnet-4-5-20250929-v1:0 \
  --input-data-config '{"s3InputDataConfig":{"s3Uri":"s3://acme-batch/in/2026-08-11/"}}' \
  --output-data-config '{"s3OutputDataConfig":{"s3Uri":"s3://acme-batch/out/2026-08-11/"}}' \
  --timeout-duration-in-hours 24

Note that the model id can be an inference profile id. AWS documents passing a profile as modelId on CreateModelInvocationJob so the job draws on compute across Regions, with output still landing in the source Region’s bucket. For a large job that is usually worth doing.

Monitoring without polling S3

GetModelInvocationJob and ListModelInvocationJobs return progress counters, so you never need to list the output prefix to find out how far along a job is: totalRecordCount, processedRecordCount, successRecordCount and errorRecordCount. Progress is processedRecordCount / totalRecordCount. AWS notes the counters read 0 before processing starts and may lag by up to a minute while the job runs, so a monitor that alerts on “no progress in 30 seconds” will page you for nothing.

Better still, do not poll at all. AWS documents EventBridge notifications for Bedrock job state changes, which turns a completed batch into an event that starts the next stage of your pipeline instead of a cron job that checks. The response also includes a modelInvocationType field telling you whether the job used the InvokeModel or Converse format — useful when you are reading somebody else’s job and need to know how to parse its output.

Reading the output

Bedrock writes one output JSONL per input JSONL. Each line carries the recordId, the modelInput echoed back, and either a modelOutput or — and this is the part to design for — an error object in place of modelOutput:

{"recordId":"3223593EFGH","modelInput":{"inputText":"Roses are red, violets are"},"modelOutput":{"inputTextTokenCount":8,"results":[{"tokenCount":3,"outputText":"blue\n","completionReason":"FINISH"}]}}
{"recordId":"1223213ABCD","modelInput":{"inputText":"Hello world"},"error":{"errorCode":400,"errorMessage":"bad request"}}

Three consequences for the consumer you write:

  • Order is not preserved. AWS states the order of records in the output file is not guaranteed to match the input. Join on recordId; never zip by line number.
  • Failure is per record. A job can complete with thousands of successes and a hundred error lines. Branch on the presence of the modelOutput key rather than assuming it, and route the failures to a retry file rather than dropping them.
  • The output shape follows the invocation type. For a Converse job, modelOutput is a Converse response body, so the text is at modelOutput.output.message.content[0].text and the stopReason values are the same nine you handle synchronously. A max_tokens stop in a batch of 50,000 is silent unless you look for it.

Alongside the output files is manifest.json.out, which is the file to read first. It contains totalRecordCount, processedRecordCount, successRecordCount, errorRecordCount, inputTokenCount and outputTokenCount for the whole job. The last two are the ones worth keeping: they are the job’s actual token consumption, which is what you multiply by the batch rate to know what the run cost, without reconstructing it from per-record responses.