Skip to content

JSON Mode Output in the DeepSeek API

8 min read · updated August 11, 2026

DeepSeek’s JSON mode makes the output parseable. It does not make the output correct, and it does not make it match your schema — those are three different problems and only the first one is solved by a parameter.

Enabling it, and its two conditions

Set response_format to an object with type: "json_object" on a request to the chat endpoint. The mechanism underneath is constrained decoding: at each generation step the sampler is restricted to tokens that can continue a valid JSON document, so a syntax error becomes structurally impossible rather than merely unlikely.

{
  "model": "deepseek-chat",
  "messages": [...],
  "response_format": {"type": "json_object"},
  "max_tokens": 1024
}

The feature is documented for the chat endpoint. Whether the reasoning endpoint accepts it has varied by model version — it was listed among the unsupported features for R1 — so verify against DeepSeek’s JSON-mode guide rather than assuming parity between the two endpoints.

The two documented conditions

DeepSeek’s documentation attaches two requirements to the mode, and both are easy to violate without noticing.

The prompt must mention JSON, and should show the shape. The word must appear in the system or user message, and the documentation asks for an example of the desired format alongside it. The first half is a guard: constrained decoding forces valid JSON out of a model that was about to write prose, and the result is a model fighting the constraint — a syntactically perfect document containing nothing useful. The second half is the part people skip, and it is what actually determines whether the keys come back as you expect.

Set a reasonable max_tokens. The documentation says this explicitly and the reason is specific to this mode: a truncated JSON document is invalid JSON. Any other truncated output is merely incomplete; here you lose the whole response, because the closing braces were never generated. See the defaults — the chat endpoint’s is smaller than most people assume.

A complete request

from openai import OpenAI
import json

client = OpenAI(api_key="sk-...", base_url="https://api.deepseek.com")

SYSTEM = """Extract the shipment details.
Reply with a JSON object in exactly this shape:

{
  "tracking_id": "string",
  "carrier": "string",
  "delivered": true,
  "items": 0
}

Use null for any field the text does not state."""

resp = client.chat.completions.create(
    model="deepseek-chat",
    messages=[
        {"role": "system", "content": SYSTEM},
        {"role": "user", "content": "Parcel 9XG-4412 went out with PostNL, three boxes, still in transit."},
    ],
    response_format={"type": "json_object"},
    max_tokens=512,
    temperature=0,
)

choice = resp.choices[0]
if choice.finish_reason == "length":
    raise RuntimeError("truncated before the JSON closed; raise max_tokens")

data = json.loads(choice.message.content)   # syntax is safe; contents are not
print(data)

Three things in that snippet are doing real work. The example object in the system message is what communicates the key names, because nothing in the request declares them. temperature=0 reduces variation in a task where creativity has no value. And the finish_reason check runs before the parse, so a truncation reports itself rather than surfacing as a confusing JSONDecodeError pointing at the last character.

What it guarantees and what it does not

  • Guaranteed: it parses. Balanced braces, quoted keys, no trailing commas, no prose wrapper, no markdown fence. The class of bug where a model returns a fenced code block containing JSON disappears.
  • Not guaranteed: your keys. The constraint is “valid JSON”, not “valid instance of your schema”. {"result": "ok"} satisfies it perfectly while satisfying you not at all.
  • Not guaranteed: your types. A field you expect as a number can arrive as a string, and a boolean can arrive as "true". Coerce and validate; do not index straight into the parsed object.
  • Not guaranteed: the values are right. Extraction errors and outright invention are unaffected by the output format. This is the one people forget, because well-formed structure reads as confidence.

The remedy for the middle two is a schema validator on your side — Pydantic, zod, or a plain JSON Schema check — with one retry that feeds the validation error back as a message. Two attempts resolve the large majority of shape failures, and a hard failure after that is more honest than a coerced guess.

Where you need the model to produce arguments matching a declared schema rather than a free-form object, function calling is a better fit: the schema is part of the request, so the model is told what shape to produce instead of being shown an example of it.

The failures that remain

  • Empty content. DeepSeek has documented that JSON mode can occasionally return an empty content, and advises adjusting the prompt. Treat empty-but-successful as a retryable outcome in code rather than as a parse failure — it will happen at some rate whatever you do.
  • Truncation. Covered above and worth repeating because it is the leading cause of parse errors in this mode. Check finish_reason first, always.
  • A valid document that is not an object. Depending on how the prompt is worded the model may return an array at the top level. If your code assumes a dict, assert the type after parsing.
  • Prompting for JSON without enabling the mode. The reverse mistake. Asking for JSON in prose and not setting response_format gets you a markdown fence eventually, and the fence-stripping code that follows is exactly what this parameter exists to delete.
  • Using stop sequences to end the object. Stopping on a closing brace removes it from the output and leaves you with invalid JSON, because the matched sequence is trimmed. Do not combine the two.

Streaming a JSON response

JSON mode and stream: true are compatible, and combining them is usually a mistake. The deltas are fragments of a JSON document split at token boundaries that have nothing to do with its structure, so a chunk is far more likely to be ing_id": "9XG- than a complete field. Nothing is parseable until the last content chunk has arrived.

That removes the reason to stream in the first place. Streaming exists so a reader sees output before generation finishes; a machine consumer gains nothing from receiving an unparseable prefix earlier, and you have added buffer management, the empty-choices usage chunk and partial-fragment handling for no benefit. If the response is going to a program, request it non-streamed.

There is one case where it is worth the trouble: a long document generated for a human who would otherwise wait thirty seconds, where a progress indication matters. Even then, treat the accumulated buffer as opaque until the stream terminates, and check finish_reason on the final chunk before parsing — the streaming page covers where that value arrives and why a stream that ends without a terminator is a failure rather than a short answer. Incremental JSON parsers exist, but they add a dependency and a class of bug in exchange for showing a user half an object.

Validating on your side

Because the mode guarantees syntax and not shape, the request is only half of a working extraction. The other half is a validator and one retry that tells the model what was wrong, which converts most shape failures into successes without a second model or a larger prompt.

from pydantic import BaseModel, ValidationError
from typing import Optional
import json

class Shipment(BaseModel):
    tracking_id: Optional[str]
    carrier: Optional[str]
    delivered: Optional[bool]
    items: Optional[int]

def extract(text: str, attempts: int = 2) -> Shipment:
    messages = [
        {"role": "system", "content": SYSTEM},
        {"role": "user", "content": text},
    ]
    last = None
    for _ in range(attempts):
        resp = client.chat.completions.create(
            model="deepseek-chat",
            messages=messages,
            response_format={"type": "json_object"},
            max_tokens=512,
            temperature=0,
        )
        choice = resp.choices[0]
        if choice.finish_reason == "length":
            raise RuntimeError("truncated; raise max_tokens")
        raw = choice.message.content or ""
        try:
            return Shipment.model_validate_json(raw)
        except ValidationError as e:
            last = e
            messages += [
                {"role": "assistant", "content": raw},
                {"role": "user", "content":
                 f"That did not validate: {e}. Return corrected JSON only."},
            ]
    raise last

Three properties of that loop are worth copying rather than the code itself. The validation error is fed back verbatim, because it names the field and the expected type more precisely than any instruction you would write by hand. The failed attempt is included as an assistant message, so the model is correcting its own output rather than starting over. And the retry count is small — a second failure after a specific correction is usually a prompt problem, and a third attempt spends money to confirm it.

Empty content deserves its own branch here rather than falling into the validation path, because the fix is different: a retry of the identical request generally resolves it, whereas feeding an empty string back as a correction gives the model nothing to correct.