Skip to content

The GPAI Training Content Summary Template

9 min read · updated August 11, 2026

Most coverage of Article 53(1)(d) stops at “you must publish a summary of your training data”. The useful question is what the document looks like, because the AI Office publishes a template and the template is what fixes the level of detail.

The duty in Article 53(1)(d)

Article 53(1)(d) of Regulation (EU) 2024/1689 requires providers of general-purpose AI models to draw up and make publicly available a sufficiently detailed summary about the content used for training the model, according to a template provided by the AI Office. The text is on EUR-Lex. Two clauses in that sentence carry all the weight: sufficiently detailed, which is a standard, and according to a template, which is a form.

The template converts a vague standard into a fillable document, which is why it matters more than the article. The AI Office published one in July 2025 alongside an explanatory notice, and both are maintained by the Commission on its digital strategy site. It is a living document: the notice contemplates revision as practice develops, so a summary drafted against an early version is not automatically compliant against a later one.

Not legal advice. The template's fields and the guidance around them are Commission material rather than statute, and they change more often than the Act does — verify the current version with the AI Office before drafting or updating a summary.

The shape of the template

The published template is organised in three parts, and the shape is more informative than any single field:

  • General information. Identification of the provider and of the model or model family the summary covers, the modalities involved, the approximate overall size of the training data, and the period over which the data was collected. This is the section that establishes what the document is about, and it is where model families get resolved: one summary may cover several released versions if they share the training corpus.
  • List of data sources. The substantive section, broken down by how the data was obtained. Covered below.
  • Relevant data-processing aspects. Measures taken to respect text-and-data-mining reservations under Article 4(3) of the CDSM Directive, and measures to identify and remove illegal content from the training data.

That third section is where Article 53(1)(c) and Article 53(1)(d) meet. The copyright policy is the control; this part of the summary is where the provider states publicly what the control does. The two obligations are distinct — see the copyright policy page — but a summary that describes opt-out handling inconsistently with the actual policy is a visible problem. The reservation mechanism the section refers to is on the Article 4 CDSM opt-out page.

The data sources section

The template does not ask for a file list. It asks the provider to describe sources by category, with the categories chosen so that a rightsholder or a regulator can locate themselves in the answer:

  • Large publicly available datasets. Named, where the dataset has a name and is identifiable — the level at which a reader can go and look at the thing.
  • Data crawled or scraped from the open web. Described with the crawlers used and, importantly, a listing of the most significant domain names by volume. This is the field that turns the summary from a genre description into something checkable: a publisher can search for its own domain.
  • Data licensed from third parties. Described by category and, where feasible, by counterparty type, acknowledging that individual contracts may be confidential.
  • User data. Data from the provider's own services or users, where used for training.
  • Synthetic data. Including whether it was generated by the provider's own models or by third-party models.
  • Other sources. Anything not captured above, so the categories do not become a way of omitting things.

The domain-listing field is where the tension in this obligation is most visible. Recital 107 frames the summary as generally comprehensive in scope rather than technically detailed, and explicitly contemplates that it should not require listing every work. Providers read that as protecting composition detail; rightsholders read the domain field as the minimum needed to exercise any right at all. Where the line falls is not settled, and it will be settled by AI Office practice and eventually by review, not by reading the article harder.

What “sufficiently detailed” is doing

Recital 107 supplies the purpose test: the summary should be sufficiently detailed to facilitate parties with legitimate interests, including copyright holders, in exercising and enforcing their rights under Union law. That is the interpretive anchor. A summary is not measured against a page count; it is measured against whether a rightsholder could use it to work out whether to act.

Two consequences follow that providers underestimate. First, aggregation that makes the answer unusable is a compliance risk even if every statement in it is true — “publicly available web data” as an entire answer defeats the stated purpose. Second, the standard is asymmetric across sections: composition proportions can be approximate, but the identification of what was crawled is the part the purpose test actually needs.

Publishing it and keeping it current

The summary must be made publicly available, which in practice means a stable page the model card or documentation points to, reachable without registration. There is no central register that hosts it for you.

Keeping it current is the part that is easy to lose. A new model version trained on additional data is a new state of affairs, and the AI Office guidance treats the summary as something to update rather than to file once. The version and date on the document are therefore substantive rather than housekeeping.

On timing: the GPAI obligations in Chapter V applied from 2 August 2025. Under Article 111(3), models placed on the market before that date have until 2 August 2027. Enforcement against model providers is the Commission's, with the Article 101 ceiling of 3% of worldwide annual turnover or EUR 15 million, whichever is higher.

Application dates here are taken from Regulation (EU) 2024/1689 as adopted. The AI Act was subsequently amended by Regulation (EU) 2026/1744, the digital omnibus on AI, published in the Official Journal on 24 July 2026 and in force from 27 July 2026. What that instrument moved was the high-risk timetable rather than the GPAI one, so the dates above are stated unchanged — but they were not verified against the amending text, and the consolidated version of the AI Act on EUR-Lex is the authority.