Skip to content

Training Data and Copyright, by Jurisdiction

12 min read · updated August 4, 2026

There is no single answer, and anyone who gives you one is describing a jurisdiction without naming it. In the EU there is a statutory exception with an opt-out. In the UK there is no commercial exception. In the US the question is fair use and it is being answered a case at a time, with the answers so far pointing in different directions on different facts.

Information, not legal advice. Reviewed 4 August 2026. This is the most actively litigated area covered in this cluster; several of the decisions described are first-instance and under appeal, and settlements have removed some questions from the courts without answering them. Read the judgments rather than the coverage. The litigation tracker page follows the cases; this page is about the statutory position each case sits inside.

The short answer, per jurisdiction

JurisdictionDescription
European UnionLawful in principle under the text and data mining exception in Article 4 of Directive (EU) 2019/790, for lawfully accessible works, unless the rightholder has reserved the right in an appropriate manner — machine-readable for content made publicly available online. A broader exception in Article 3 covers research organisations and cannot be opted out of.
United KingdomThe text and data mining exception in section 29A of the Copyright, Designs and Patents Act 1988 covers non-commercial research only. There is no commercial TDM exception. Reform has been consulted on and has not been enacted.
United StatesNo statutory exception. The question is fair use under 17 U.S.C. § 107, assessed on the facts. Rulings in 2025 went different ways on different facts, and the source of the copies has emerged as decisive.
JapanArticle 30-4 of the Copyright Act permits exploitation for information analysis where the purpose is not enjoyment of the expression, subject to a proviso for uses that would unreasonably prejudice the rightholder. Broad, but narrower than its reputation.
SingaporeA computational data analysis exception in the Copyright Act 2021 permits copying for analysis subject to lawful access, and cannot be excluded by contract. This page does not give the section number.
ChinaNo general TDM exception. The generative AI measures require training data to come from lawful sources and not to infringe intellectual property rights, which makes lawfulness of acquisition a regulatory as well as a private-law question.

EU: a statutory exception with an opt-out

The Copyright in the Digital Single Market Directive created two text and data mining exceptions. Article 3 permits reproductions and extractions by research organisations and cultural heritage institutions for scientific research on works they have lawful access to, and it cannot be overridden by contract or by a rightholder reservation. Article 4 is the general one: it permits reproductions and extractions of lawfully accessible works for TDM, but only where the use has not been expressly reserved by the rightholder in an appropriate manner — and for content made publicly available online, that means machine-readable means.

Three consequences fall out of that structure and each is contested at the edges.

  • “Appropriate manner” is undefined. A robots.txt directive, a term in a website’s conditions of use, a metadata field in an image: which of these is an effective reservation is a real question with real money on it, and there is no settled answer.
  • The research exception has been read broadly at first instance. The Hamburg regional court held in September 2024, in a claim by a photographer against LAION, that assembling a dataset of image links for research fell within the German implementation of the research TDM exception. It is a first-instance decision about dataset creation rather than about model training, and it was appealed.
  • The AI Act adds a compliance duty on top. Providers of general-purpose models must have a policy to identify and respect Article 4(3) reservations, and must publish a summary of training content. Recital material asserts this applies to models placed on the Union market wherever training occurred, which is a deliberate attempt at extraterritorial reach and one of the more likely candidates for a future challenge. See the general-purpose model obligations.

Separately, and importantly, the TDM exceptions are about the copying done to train. They say nothing about a model that reproduces a work in its output. A Munich regional court ruled in November 2025 in a claim brought by the German collecting society GEMA that a model reproducing song lyrics in its outputs infringed, on a memorisation theory. That is a different question from whether training was lawful, and it is the question that is now moving.

UK: no commercial exception, and a stalled reform

Section 29A of the CDPA permits copies for computational analysis for the sole purpose of research for a non-commercial purpose, by a person with lawful access. Commercial training is not covered. There is no general fair use doctrine in UK law to fall back on, and the fair dealing exceptions are closed categories that do not fit.

The position has been under review for years without resolution. A 2021 consultation proposed a broad commercial TDM exception, which was announced and then abandoned after opposition from rightholders. A further consultation ran from December 2024 into 2025 proposing an exception with a rightholder opt-out plus transparency requirements. During the passage of the Data (Use and Access) Act 2025 there was a protracted fight over amendments that would have imposed transparency obligations on AI developers about the works they had used; those amendments did not survive into the Act, and the government committed instead to reports on the subject. As at this review date, the UK position is unchanged: no commercial TDM exception.

The most significant UK case, Getty Images against Stability AI, is less useful than its prominence suggests. Getty abandoned its principal training and output copyright claims during the trial in 2025, largely because the training happened outside the United Kingdom, and the judgment that followed in November 2025 dealt with trade marks and with a secondary infringement argument about whether model weights are an infringing copy. The UK courts have therefore still not ruled on whether training on copyright works in the UK infringes.

US: fair use, decided case by case

There is no TDM exception in US law. Everything turns on the four fair use factors in section 107, and 2025 produced the first substantive rulings.

DecisionDescription
Thomson Reuters v. Ross Intelligence (D. Del., February 2025)Fair use rejected. The defendant used Westlaw headnotes to train a non-generative legal research tool that competed directly with the plaintiff. The competitive substitution was the decisive factor. Certified for interlocutory appeal.
Bartz v. Anthropic (N.D. Cal., June 2025)Training on books to build a language model was held to be exceptionally transformative and fair. But retaining a library of pirated copies was held not to be fair use — the same use of the same works came out differently depending on how the copies were obtained. The case subsequently settled on a large class-wide basis, so the ruling was never tested on appeal.
Kadrey v. Meta (N.D. Cal., June 2025)Summary judgment for the defendant on fair use, but the opinion is explicit that this is not a holding that training on copyrighted works is generally lawful — it is a holding that these plaintiffs failed to develop a record on market dilution. The judge effectively wrote a roadmap for the next plaintiff.
New York Times v. OpenAI and Microsoft (S.D.N.Y.)Motions to dismiss were largely denied in 2025 and the case proceeded into discovery, consolidated with related actions. Live issues include output regurgitation and the retention of user conversation logs.

Three things are worth extracting from that set. The source of the copies matters independently of the use: acquiring works unlawfully is a separate wrong from training on them. Market substitution is doing most of the work in the fourth factor, and the emerging theory that generative output dilutes the market for the class of works, rather than substituting for any particular one, is untested at appellate level. And settlements are removing questions from the courts: a very large settlement resolves the parties’ dispute and creates a price signal, but it does not create precedent.

The US Copyright Office published a pre-publication part of its generative AI report on training in May 2025. It concluded, in substance, that transformativeness varies with the use and that some uses will be fair and others will not — consistent with the case law and not a rule anyone can plan around.

Japan and Singapore: the permissive end

Japan’s Article 30-4 permits exploitation of a work where the purpose is not to enjoy, or cause another to enjoy, the thoughts or sentiments expressed in it — information analysis being the named example. This is broader than the EU exception because it has no opt-out.

It is narrower than its reputation, though, and the Agency for Cultural Affairs said so in 2024. The proviso excludes uses that would unreasonably prejudice the copyright owner’s interests; training intended to produce outputs in the style of and substituting for a specific creator’s work, or training on a database sold for the purpose of information analysis, may fall outside. And the exception covers the training, not the output: a generated work that reproduces a protected expression can infringe in the ordinary way.

Singapore’s computational data analysis exception, introduced in its 2021 Copyright Act, permits copying for analysis where the user had lawful access to the material, and cannot be excluded by contract. The lawful access condition is the operative limit.

China

China has no TDM exception. Its fair use provisions are a closed list that does not comfortably accommodate training. What it has instead is a regulatory requirement, in the generative AI measures, that training data come from lawful sources and not infringe intellectual property rights — which converts the question from a private dispute into a condition of operating the service. Enforcement of that condition runs through the filing process described on the China regulation page, not through the courts.

What is genuinely unsettled, and what turns on it

The honest summary is that four questions are open and each would change the economics of model development if answered one way rather than the other.

  • Is market dilution a cognisable harm under the fourth fair use factor? If a court accepts that flooding a market with machine-generated substitutes for a genre harms the market for the works the model learned from, the fair use defence narrows sharply even for lawfully acquired training data. If it does not, the defence holds for anyone who paid for their copies.
  • What counts as an effective reservation of rights under EU Article 4(3)? If a term in website conditions suffices, almost all commercial web content is reserved and the exception becomes narrow. If only a machine-readable protocol counts, the burden sits on rightholders to adopt one.
  • Do model weights contain copies? The secondary infringement argument rejected at first instance in the UK, and the memorisation theory accepted at first instance in Munich, are two sides of this. If weights can be an infringing article, distributing a model becomes a distinct act of infringement, which changes what open-weights release means legally.
  • Does extraterritorial reach hold? The AI Act’s copyright policy duty purports to apply to training done anywhere, for models placed on the EU market. Whether that survives challenge determines whether jurisdiction shopping for training location works.

What a company building on models should do

If you are not training foundation models, most of the above is background. Four things do reach you.

  1. Check the indemnity, and its conditions. Major model providers offer IP indemnities for outputs. They are conditional — typically on using the service as intended, not disabling safety features, not supplying infringing material as input and not deliberately provoking reproduction. The conditions are where the protection is won or lost, and they belong in your contract review.
  2. Keep the training content summary you were given, with the date you retrieved it. It is your evidence of what you knew and relied on.
  3. Do not fine-tune on scraped data casually. If you fine-tune, you are doing the copying, and the analysis on this page applies to you directly rather than to your vendor. Record provenance for every dataset; provenance tracking is cheap before the fact and impossible after it.
  4. Treat verbatim reproduction as a bug. Whatever the law decides about training, output that reproduces a substantial part of a protected work is a problem in every jurisdiction on this page. Detecting it is a product requirement, not a legal one.