Skip to content

Edge Functions on Netlify for Streaming a Model Response

10 min read · updated August 11, 2026

A Netlify Edge Function is Deno running in an isolate with a 50 millisecond CPU budget per request. That budget sounds fatal for streaming a model response until you notice what it excludes: time spent waiting on a fetch is not CPU time. Forwarding a provider’s stream costs almost nothing. Parsing every token does not.

The runtime you are writing for

Netlify documents Edge Functions as running in a Deno runtime with broad Web API support: fetch, Request, Response, the Streams API (ReadableStream, WritableStream, TransformStream), Web Crypto, TextEncoder and TextDecoder with their stream variants, WebSocket and URLPattern. Node built-ins are available with the node: prefix, Deno modules by URL import, and npm packages by name — with the caveat Netlify states plainly, that npm support is in beta and packages relying on native binaries or dynamic runtime imports may not work.

That caveat is why an SDK that works in a Node function can fail here. For streaming specifically you need none of it: the provider’s HTTP endpoint plus fetch is the whole dependency list.

Netlify’s documented edge limits, read on 11 August 2026: 20 MB code size after compression, 512 MB memory per set of deployed edge functions, 50 ms CPU execution time per request, and a 40 s response header timeout.

From Netlify’s Edge Functions limits page. Isolate CPU budgets are exactly the kind of figure a platform revises; confirm before designing around the margin.

The 50 ms rule shapes the code

Netlify is explicit that the CPU limit counts only time spent running your scripts, not time waiting for resources or responses. This single distinction decides the architecture:

  • Cheap: calling the provider, returning upstream.body as the response body, setting headers. The isolate is idle for almost the entire request.
  • Expensive: decoding every chunk to text, splitting on newlines, JSON.parse-ing each SSE data: frame, re-encoding, accumulating the full answer in a string. A long completion is thousands of small chunks, and each one costs a slice of the same 50 ms.

The 40 second response header timeout is the other shaping constraint, and it is a deadline on the first byte of the response, not on the whole exchange. Return the Response as soon as the upstream headers arrive and let the body drain afterwards; do not await the completion and then construct a response.

Building the function

  1. Create the file at netlify/edge-functions/chat.ts and declare its route with a config export. Netlify’s declaration reference documents path as a URLPattern expression that must start with /, plus optional excludedPath, pattern, method, header, onError and cache.
  2. Call the provider with stream: true and check both upstream.ok and the presence of a body before handing it on. A failed upstream call still has a body — an error JSON — and piping that through as if it were a token stream produces a client that renders an error message one character at a time.
  3. Return the upstream body directly. No transform, no decode:
// netlify/edge-functions/chat.ts
import type { Config } from "@netlify/edge-functions";

export default async (request: Request) => {
  const { prompt } = await request.json();

  const upstream = await fetch("https://api.openai.com/v1/chat/completions", {
    method: "POST",
    headers: {
      "Content-Type": "application/json",
      Authorization: "Bearer " + Netlify.env.get("OPENAI_API_KEY"),
    },
    body: JSON.stringify({
      model: "gpt-4o-mini",
      stream: true,
      messages: [{ role: "user", content: prompt }],
    }),
  });

  if (!upstream.ok || !upstream.body) {
    return new Response(
      JSON.stringify({ error: "upstream", status: upstream.status }),
      { status: 502, headers: { "Content-Type": "application/json" } },
    );
  }

  return new Response(upstream.body, {
    headers: {
      "Content-Type": "text/event-stream; charset=utf-8",
      "Cache-Control": "no-cache, no-transform",
      Connection: "keep-alive",
    },
  });
};

export const config: Config = { path: "/api/chat", method: "POST" };

Three header choices are doing real work here. Content-Type: text/event-stream tells every intermediary that this is an event stream rather than a document to be collected. no-transform in Cache-Control asks intermediaries not to recompress or otherwise rewrite the body, which is the class of behaviour that turns a stream into one late chunk. And no-cache keeps a CDN from serving one user’s completion to the next.

If you do need to transform — stripping the SSE envelope so the client receives plain text, say — do it in a TransformStream and keep the per-chunk work to a substring operation. Every JSON.parse per token is a bill against the same 50 ms.

Failing after the headers are sent

This is the part streaming tutorials leave out and production finds immediately. Once you have returned the Response, the status code is committed. HTTP has no facility to retract a 200 and replace it with a 502 because something went wrong at token fifty. Whatever the provider does from that point — a rate limit mid-generation, a dropped connection, a content filter terminating the stream — your client has already been told the request succeeded.

A pass-through function makes this sharper, because it deliberately does not inspect the bytes. If the upstream stream ends early, the browser sees a body that simply stops, which is indistinguishable from a completed short answer. The consequences are worth being explicit about:

  • Do all the checking you can before returning. The upstream.ok and upstream.body test in the code above is the last moment at which a real status code is available to you. Everything after it is in-band.
  • Give the client a completion signal it can check. The provider’s own terminator — the [DONE] sentinel in an OpenAI-shaped SSE stream — is the thing to look for. A client that renders text until the socket closes cannot tell truncation from completion; a client that requires the terminator can, and can retry or mark the answer partial.
  • Do not retry a stream by restarting it silently. The user has already seen fifty tokens. Re-running the request produces a different continuation and either duplicates or contradicts what is on screen, and it bills you twice. Surface the partial state instead.
  • The 40-second header deadline protects you here. Because it applies to the first byte, a provider that hangs before producing anything fails while you can still return a real error. It is only after that first byte that you lose the ability, which is another argument for returning as soon as the upstream headers arrive rather than waiting for content you could have forwarded.

Proving it streams

The whole failure mode of this page is a function that looks correct and delivers everything at once, so verify rather than assume. Run netlify dev locally, then against the deployed URL use curl -N, which disables curl’s own output buffering:

curl -N -X POST https://your-site.netlify.app/api/chat \
  -H 'Content-Type: application/json' \
  -d '{"prompt":"Count slowly from one to twenty."}'

Text appearing progressively is a working stream. Text appearing all at once after a pause is not, and the next thing to check is whether anything between the function and you is collecting the body — a proxy, a compression layer, or the client. Add -w '\ntotal:%{time_total} first_byte:%{time_starttransfer}\n' to see the two numbers separately: a healthy stream has a small time-to-first-byte and a much larger total.

Reading it in the browser

A correct server stream is still invisible if the client awaits the whole body. await response.text() and await response.json() both resolve only when the body ends. Read the reader:

const response = await fetch("/api/chat", {
  method: "POST",
  headers: { "Content-Type": "application/json" },
  body: JSON.stringify({ prompt }),
});

const reader = response.body!.pipeThrough(new TextDecoderStream()).getReader();
for (;;) {
  const { value, done } = await reader.read();
  if (done) break;
  for (const line of value.split("\n")) {
    if (!line.startsWith("data: ")) continue;
    const payload = line.slice(6);
    if (payload === "[DONE]") continue;
    render(JSON.parse(payload));
  }
}

Parsing on the client rather than in the edge function is deliberate: the browser has no 50 ms budget, and moving the per-token work there costs you nothing.