Skip to content

Choosing Between a Nested and a Flat Schema for Document Extraction

10 min read · updated August 11, 2026

Flat schemas are easier to store, easier to diff and easier to show in a table. Nested schemas describe the document honestly. The choice is usually presented as taste; it is not, and the deciding question is a narrow one about how many times a thing can occur.

The same invoice, both ways

Take an invoice with a header, two line items and two tax rates. Flat, it looks like this — and the awkwardness is visible immediately.

invoice_number        = "INV-4417"
invoice_date          = "2026-03-04"
supplier_name         = "Northgate Components Ltd"
line_1_description    = "Bearing housing, 40mm"
line_1_quantity       = 12
line_1_unit_price     = 3450        # minor units
line_2_description    = "Freight"
line_2_quantity       = 1
line_2_unit_price     = 9500
tax_1_rate            = 20.0
tax_1_amount          = 10190
tax_2_rate            = 0.0
tax_2_amount          = 0
total_amount          = 61590

Three problems are already present. The schema has a fixed maximum number of lines, and the invoice with nineteen lines does not fit. The index in line_2 means nothing — it is the order the lines happened to be printed in, and it will differ if the same goods are invoiced again. And the relationship between tax_1_rate and tax_1_amount is expressed only by a shared prefix, which no validator enforces.

{
  "invoice_number": "INV-4417",
  "invoice_date": "2026-03-04",
  "supplier": { "name": "Northgate Components Ltd" },
  "lines": [
    { "description": "Bearing housing, 40mm", "quantity": 12, "unit_price_minor": 3450 },
    { "description": "Freight",               "quantity": 1,  "unit_price_minor": 9500 }
  ],
  "taxes": [
    { "rate_percent": 20.0, "amount_minor": 10190 },
    { "rate_percent": 0.0,  "amount_minor": 0 }
  ],
  "total_amount_minor": 61590,
  "currency": "GBP"
}

The nested form has none of those problems and introduces its own: it cannot be a row in a table, a consumer has to walk it, and a change to one line item is a change to a document rather than to a column.

The rule: cardinality decides

The rule is short. Nest where the cardinality is unbounded or unknown. Flatten where it is exactly one.

An invoice has exactly one invoice number, so invoice_number is a scalar and nesting it inside a header object buys nothing but a longer path. It has an unknown number of lines, so lines are an array of objects. It has exactly one supplier, but the supplier has a name, an address and a tax registration, so it is an object — not because suppliers repeat, but because grouping keeps the three fields together when one of them is corrected.

The test that resolves most arguments: could this field ever need to appear twice on one document, for reasons the document type allows? If yes, it is an array, today, even though every document you have seen has one. That is the same argument made from the other direction on designing for variants you have not seen.

Field paths are the real reason

Here is the consideration that usually settles it, and it is rarely mentioned. Everything an extraction pipeline does after extraction is keyed on a field identifier: the confidence score, the reviewer’s correction, the highlight rectangle on the page image, the audit row recording who changed what. That identifier has to be stable, unique, and meaningful to a human reading a log eighteen months later.

A nested schema gives you one for free in the form of a path, written in the JSON Pointer syntax standardised as RFC 6901 or in the equivalent dotted form:

/lines/1/unit_price_minor
/taxes/0/rate_percent
/supplier/name

A flat schema gives you a name that encodes the same thing in a string you have to parse: line_2_unit_price means the same as the first path, but a consumer that wants “every unit price” has to match a pattern rather than walk a structure, and a consumer that wants “the second line” has to know your prefix convention.

There is one genuine advantage on the flat side here, and it is worth naming: array indices are positional, so if a re-extraction reorders the lines, /lines/1 now refers to a different line, and a correction recorded against that path is attached to the wrong thing. The fix is not to flatten but to give array elements a stable identity — a hash of the source region, or an explicit line identifier read off the document if it has one — and to key corrections on that. This matters most for correcting one field without re-running the extraction and for per-field confidence, both of which are meaningless without a path that survives.

What each shape costs downstream

  • Relational storage. Flat is one wide row. Nested is a parent row plus a child table per array, which is more schema but is also the shape any reporting query wants — summing tax by rate across a year is trivial against a child table and painful against numbered columns.
  • Correction. Updating one flat field is a column update. Updating one nested field is a document rewrite unless you store arrays as rows, which is a strong argument for doing so.
  • Export. Anything that has to become a spreadsheet gets flattened eventually. Flattening a nested structure at export time is mechanical; unflattening a flat structure back into arrays requires knowing the prefix convention, so the nested form is the better thing to store even when the deliverable is flat.
  • Diffing two extractions. Flat diffs are trivially readable. Nested diffs need a path-aware comparison, and array reordering shows up as a large spurious diff unless you match on element identity first.

What each shape costs the model

The schema is part of the request, so its shape has a cost and an effect on behaviour. Three points are worth knowing before choosing a deep structure.

A large schema consumes input tokens on every call, and a deeply nested one with repeated object definitions consumes many. Where the same object shape appears in several places, define it once and reference it rather than inlining it, if the provider’s constrained mode supports references. Second, providers impose limits on schema size, nesting depth and the number of enum values, and those limits are the kind of thing that bites at the worst moment; check them before designing a five-level structure. Third, a required nested object that the document does not contain has to be filled with something, and a model asked for a mandatory guarantor object on a document with no guarantor will produce one with null fields — or, worse, a plausible one. Make the object itself nullable rather than making its fields nullable.

The shape most extractions end up with

The stable answer is neither. Flatten the header, nest the repeats, and keep the nesting shallow — two levels is almost always enough for a business document.

That means scalar fields at the top for everything with cardinality one, one object per real-world entity that has several attributes, one array per group that repeats, and no level that exists only for tidiness. A header object wrapping fields that are already unambiguous at the top level is the most common piece of gratuitous nesting, and every path in the document gets longer to pay for it. For the general design principles behind schemas a model has to fill, see designing a schema for AI.