Streaming a Model Response From a Vercel Edge Function
10 min read · updated August 11, 2026
A streamed response starts arriving while the model is still writing. On Vercel’s Edge runtime that is not only a perceived-speed improvement — it is what keeps a long generation inside the platform’s first-byte deadline.
Why streaming is the fix, not a nicety
Vercel documents Edge runtime functions as needing to begin sending a response within 25 seconds, after which streaming may continue for up to 300 seconds. A buffered response begins when generation ends, so a long answer puts the whole generation inside the 25-second window. A streamed response begins at the model’s first token, which typically arrives in a second or two regardless of how long the answer will eventually be.
So streaming moves a route from “must finish quickly” to “must start quickly”, and the second is a far easier promise to keep. The perceived-speed benefit is real too, and secondary to this.
The handler
The high-level route is the AI SDK, which Vercel recommends for streaming from model providers and which handles the provider’s event format for you:
// app/api/chat/route.ts
import { streamText } from "ai";
export async function POST(req: Request) {
const { prompt } = await req.json();
const result = streamText({
model: "openai/gpt-4o-mini",
messages: [{ role: "user", content: prompt }],
});
return result.toTextStreamResponse({
headers: { "Content-Type": "text/event-stream" },
});
}The lower-level route is worth writing once even if you use the SDK, because it shows where every byte comes from and it is what you will debug when something buffers. You read the provider’s stream, transform it, and enqueue into a ReadableStream of your own:
export const config = { runtime: "edge" };
export default async function handler(req: Request): Promise<Response> {
const { prompt } = await req.json();
const upstream = await fetch("https://api.openai.com/v1/chat/completions", {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.OPENAI_API_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
model: "gpt-4.1-mini",
stream: true,
messages: [{ role: "user", content: prompt }],
}),
});
if (!upstream.ok || !upstream.body) {
return new Response("upstream failed", { status: 502 });
}
const decoder = new TextDecoder();
const encoder = new TextEncoder();
let buffered = "";
const stream = new ReadableStream<Uint8Array>({
async start(controller) {
const reader = upstream.body!.getReader();
try {
for (;;) {
const { done, value } = await reader.read();
if (done) break;
buffered += decoder.decode(value, { stream: true });
const lines = buffered.split("\n");
buffered = lines.pop() ?? ""; // keep the partial line
for (const line of lines) {
if (!line.startsWith("data: ")) continue;
const payload = line.slice(6).trim();
if (payload === "[DONE]") continue;
const delta = JSON.parse(payload)?.choices?.[0]?.delta?.content;
if (delta) controller.enqueue(encoder.encode(delta));
}
}
} catch (err) {
controller.error(err);
return;
}
controller.close();
},
});
return new Response(stream, {
headers: {
"Content-Type": "text/plain; charset=utf-8",
"Cache-Control": "no-cache, no-transform",
"X-Accel-Buffering": "no",
},
});
}The buffered variable is the part that is wrong in most hand-written versions. A network chunk is not a line: one read can deliver half an event, or three events and a fragment. Keeping the trailing partial line and prepending it to the next read is what stops an intermittent JSON.parse failure that only appears under real network conditions.
Reading it in the browser
fetch gives you a ReadableStream on response.body. Calling await response.text() instead waits for the whole thing and throws away every benefit above — it is the single most common way a working stream is rendered as a buffered one.
const res = await fetch("/api/chat", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ prompt }),
});
if (!res.body) throw new Error("no stream");
const reader = res.body.getReader();
const decoder = new TextDecoder();
for (;;) {
const { done, value } = await reader.read();
if (done) break;
const text = decoder.decode(value, { stream: true });
setOutput((prev) => prev + text); // paint as it arrives
}{ stream: true } on the decoder matters for the same reason the server-side buffer does: a multi-byte character can be split across two chunks, and without it you get a replacement character in the middle of a word for any non-ASCII text.
Errors and aborts after the first byte
Streaming changes what a failure can look like, and this is the part that separates a demo from something you can leave running. Once you have written a single byte, the status line and the headers are gone — they were sent when the stream opened. There is no way to turn a 200 into a 502 halfway through, so an upstream that dies at token 300 cannot be reported the way you normally report an error.
Calling controller.error() does the only thing available at that point: it tears the connection down mid-body. The client sees a truncated response and a network-level failure, which is indistinguishable from the user’s wifi dropping. If you want the client to be able to tell the difference, the error has to travel in-band, as content:
try {
// ... read upstream, enqueue deltas ...
} catch (err) {
// The status code is long gone. Say it in the body instead.
controller.enqueue(encoder.encode("\n\n[[stream-error]]"));
controller.close();
return;
}A sentinel is crude and it is honest: the client can look for it, render a retry affordance, and avoid saving a half-answer as if it were complete. If you are using a structured event format, the same idea is an event: error frame, which is tidier and works identically.
The other direction matters more for your bill. When the user closes the tab or navigates away, the read side of your response goes away and the cancel method on your ReadableStream runs. If you do nothing there, the upstream request is still open and the provider keeps generating tokens that nobody will ever see — and you are charged for every one of them. Forward the cancellation:
const upstreamAbort = new AbortController();
const upstream = await fetch(providerUrl, {
method: "POST",
headers,
body,
signal: upstreamAbort.signal,
});
const stream = new ReadableStream<Uint8Array>({
async start(controller) { /* ... as before ... */ },
cancel(reason) {
// Client went away. Stop paying for tokens nobody will read.
upstreamAbort.abort(reason);
},
});The incoming Request also carries a signal, and passing req.signal straight through to the upstream fetch covers the same case with less code. Use one or the other; using neither is the default, and the default is a silent cost on every abandoned generation — which, on a chat interface where people rephrase mid-answer, is not a rare event.
Proving it is actually incremental
- curl with no buffering.
curl -N -X POST https://your-app.vercel.app/api/chat -d '...'—-Ndisables curl’s own buffering. Text should appear progressively. If it appears all at once, the problem is on the server or in front of it, not in your browser code. - Timestamp the chunks. Pipe through a stamper so you can see the arrival times rather than trusting your eyes:
curl -N ... | while IFS= read -r -n1 c; do printf '%s' "$c"; done. Any per-chunk timestamping works; the point is a spread of times rather than one. - Check the deployed URL, not the dev server. Local development and production differ in exactly the layer that buffers. A stream that works on
localhostand not on a deployment is the normal shape of this bug, not a surprise. - Confirm the first byte is early. Log a timestamp on your first
controller.enqueue. Comfortably under the 25-second deadline is the property you actually care about.
When it arrives all at once anyway
A correct stream can still be delivered as one lump, and there are only a few causes worth checking:
- Something awaited the whole upstream response.
await upstream.json()orawait upstream.text()anywhere in the handler collapses the stream before you ever enqueue. This is the most common cause and it is invisible in a diff. - Compression or transformation in the path.
Cache-Control: no-transformandX-Accel-Buffering: noare the two headers that tell intermediaries to leave the body alone. They are cheap insurance even where they are not strictly required. - The client used
.text(). Covered above, and worth re-checking, because a library wrapper may be doing it on your behalf. - Chunks too small to leave the buffer. A stream that emits one or two characters at a time can sit in a socket buffer. Emitting whole words or flushing periodically usually resolves it.
- A framework response wrapper. Returning the stream through a helper that serialises the body defeats it. Return a plain
Responsewhose body is the stream.
If none of those is it, the dedicated buffering fix page goes further. And if the reason you are here is a 504 rather than a rendering complaint, the timeout page explains which deadline you crossed.