Vercel Fluid Compute for AI Workloads
10 min read · updated August 11, 2026
A function that calls a model spends most of its life doing nothing at all — suspended at an await, holding memory, using no CPU. Classic serverless billing charges for that wall-clock time. Fluid compute splits the meter, and the split is the whole story for AI workloads.
The problem with billing a waiting function
Under a duration-based model, a request that waits eight seconds for a completion and computes for fifty milliseconds is billed as eight seconds of function time. The economics push you towards not calling models from functions at all — towards a long-lived server whose idle waits are free because you already paid for the box.
That is the pressure Fluid compute is designed to remove, and Vercel is direct about the target: its documentation calls optimised concurrency especially valuable for AI applications, where tasks like fetching embeddings, querying vector databases, or calling external APIs are I/O-bound.
Three meters, not one
Vercel documents three billed resources, and they behave differently:
- Active CPU. The CPU time your code actively consumes, billed per CPU-hour. Vercel documents that you are billed only during actual code execution and not during I/O operations such as database queries or AI model calls, and that billing pauses when your code is waiting for external services.
- Provisioned Memory. The memory allocated to your instances, billed in GB-hours for the entire instance lifetime. Vercel documents that this continues billing during I/O operations and until the last in-flight request completes.
- Invocations. Counted per incoming request, regardless of success or failure.
The second bullet is the one that gets skipped and it is the one that costs money. Waiting is free for CPU and not free for memory. A function that waits is still holding its allocation, and the meter on that allocation runs.
Vercel documents the Hobby allowances as 4 hours of Active CPU, 360 GB-hours of Provisioned Memory and 1 million invocations included, with Pro billed on demand at rates that vary by region.
What it costs to wait
The published regional rates make this concrete. Vercel documents Washington D.C. (iad1) at $0.128 per Active CPU hour and $0.0106 per GB-hour of provisioned memory, and São Paulo (gru1) at $0.221 and $0.0183. Its own worked example uses a 4 GB function in São Paulo with 4 seconds of active CPU on an instance alive for 10 seconds, and arrives at $0.0002456 for CPU plus $0.0002033 for memory, $0.0004489 per invocation.
Now take the shape this page is about — a mostly-waiting model call — and follow the same arithmetic in iad1 at 2 GB, assuming 100 ms of active CPU and 8 seconds of instance lifetime:
CPU: (0.1 s / 3600) x $0.128 = $0.0000036 Memory: (2 GB x 8 s / 3600) x $0.0106 = $0.0000471 Total = $0.0000507 per invocation Assumptions: iad1 published rates, 2 GB configured memory, 100 ms active CPU, 8 s instance lifetime, one request per instance. Rates as published at the time of writing; verify before you rely on it.
Two conclusions fall out. First, memory is roughly 93% of the cost of this request — the CPU term is nearly a rounding error, which is exactly the claim Active CPU billing makes. Second, and less obvious: because memory is billed per instance rather than per request, the concurrency behaviour in the next section is what actually determines your bill. Two requests sharing one instance for those 8 seconds approximately halve the memory cost of each.
The assumptions line above is not decoration. Every number in that block is either published by Vercel or something you must measure for your own function, and a per-invocation cost derived from someone else’s memory setting is not your cost.
Instances are shared now
Vercel documents Fluid compute as allowing multiple invocations to share a single function instance, with optimised concurrency available on the Node.js and Python runtimes, and describes it as using a different isolation approach from a microVM per invocation: multiple invocations can share the same physical instance and global state concurrently.
That is the source of the cost saving and it changes a correctness assumption most serverless code was written under. Module scope is no longer per-request:
// DANGEROUS under in-function concurrency.
let currentUser: User | null = null;
export async function POST(req: Request) {
currentUser = await authenticate(req); // request B overwrites request A
const answer = await callModel(req, currentUser);
return Response.json({ answer, user: currentUser!.id }); // may be B's user
}Module-scope mutable state shared between concurrent requests is a cross-user data leak, and the window for it is exactly the seconds spent awaiting a model — the longest await in the function. Keep per-request state in the request scope, and reserve module scope for things that are genuinely shared and immutable: a client object, a compiled schema, a config constant.
Vercel also documents error isolation for this model: when an uncaught exception or unhandled rejection happens in Node.js, Fluid compute logs the error and lets current requests finish before stopping the process, so one broken request does not crash the others on that instance. Worth knowing, and not a substitute for handling your errors — the process still stops afterwards.
What actually changes in your code
- Waiting in the function stops being the expensive option. Retries, a fallback to a second provider, and a short backoff all cost almost no CPU. They still cost provisioned memory for the duration, so a 60-second retry ladder is not free — it is just far cheaper than it used to be.
- Configured memory became a bigger lever than CPU. Since memory dominates the bill for an I/O-bound function, an over-provisioned memory setting is a permanent multiplier on every request. Vercel documents defaults and maxima of 2 GB / 1 vCPU on Hobby and up to 4 GB / 2 vCPU on Pro and Enterprise.
maxDurationis a cost control, not just a safety net. Memory bills until the last in-flight request completes, so a hung upstream call holds an instance for whatever duration you allowed. Set it to something close to your real p99, not to the plan maximum.- Background work has a proper home. Vercel documents
waitUntilfor continuing execution after the response — logging, analytics, writing a trace — which keeps that work off the user’s critical path while staying inside the same instance. - Cold starts are mitigated, not eliminated. Vercel documents bytecode caching on Node.js 20 and later, applied to production environments only and not to development or preview deployments, plus function pre-warming on production deployments. A preview deployment that feels slower than production is expected behaviour rather than a regression.
Fluid compute is documented as enabled by default for new projects as of 23 April 2025, and can be set explicitly with "fluid": true in vercel.json for a specific environment or deployment. If you are on an older project and have never checked, that toggle is worth finding — for a function that mostly waits on a model API, it is the difference between paying for eight seconds of compute and paying for a tenth of a second of it.