Skip to content

Training Data Copyright: The Shape of the Dispute

5 min read · updated August 3, 2026

Whether training a model on copyrighted material without a licence is lawful is being litigated and legislated in several places at once, with different rules and different stages of resolution. The useful thing to hold is not a list of cases but the structure of the argument, because that structure is what makes each new development legible.

This is engineering and product orientation, not legal advice, and it is deliberately not a summary of current outcomes. This page carries a hard review date: if the date at the top is more than six months old, treat the framing as stale and check the current position before relying on anything here.

Why this is not a case tracker

A page listing active cases and their status is wrong within weeks, and wrong in the most damaging possible way — a stale claim about a legal outcome reads as authoritative and gets repeated. Rulings are appealed, settlements are confidential, procedural decisions get reported as substantive ones, and the same underlying question receives different answers in different courts simultaneously.

What does not go stale is the anatomy. The claims are a short list, the defences are a short list, and almost every headline you will read is one of them advancing or retreating somewhere.

The claims being made

  • Copying during dataset assembly. Acquiring and storing works to build a corpus involves reproduction, separate from anything the model later outputs. This claim is about the act of copying itself, and it is the one least dependent on what the model does.
  • Copying during training. Whether the training process makes reproductions that matter legally, and whether the resulting weights are themselves a form of copy or a non-infringing abstraction.
  • Infringing outputs. That the model produces material substantially similar to protected works. Factually narrower than the other claims — it depends on memorisation of specific content — and, when it succeeds, it is the claim with the most direct implications for users of a model rather than its builder.
  • Removal of rights-management information. That stripping attribution and licensing metadata during dataset preparation is a distinct wrong under provisions that protect such information.
  • Market harm. That the outputs substitute for the original works or displace a licensing market that would otherwise exist. Formally part of some defences rather than a standalone claim, but it is doing a great deal of the work in practice.

The defences and exceptions

The other side of the argument is not a single doctrine, and this is where jurisdiction matters most.

In the United States the central question is fair use, assessed on a multi-factor analysis in which the transformative character of the use and the effect on the market for the original tend to dominate the discussion. The analysis is fact-specific by design, which is exactly why outcomes vary between cases that sound similar in a headline.

In the European Union the relevant provisions are text-and-data-mining exceptions, one of which is broader but allows rightsholders to reserve their rights — an opt-out mechanism whose practical operation, including what counts as an effective reservation and how it interacts with models trained elsewhere, is a live question. Several other jurisdictions have their own mining exceptions with materially different scope, and a few have moved to clarify or narrow them recently. Read the current text rather than a description of it, including this one.

Alongside the litigation, licensing markets have been forming: content owners and model developers striking agreements is itself an argument in the disputes, since the existence of a licensing market bears on the market-harm analysis.

Why the answer differs by jurisdiction

Three structural reasons, worth knowing so that a result in one place is not over-read. Copyright is territorial, so a decision binds where it was made. The exceptions are drafted differently — an open-ended standard versus enumerated exceptions produces different reasoning from the same facts. And the acts happen in different places: dataset assembly, training and inference can each occur under a different law, which is one reason the location of a training run is not merely an infrastructure decision.

What changes for you, in each direction

For someone building on models rather than training them, the practical exposure is narrower than the volume of coverage suggests, and it is worth being specific about what would actually move.

  • If the broad claims against training succeed somewhere significant: expect licensing costs to appear in model pricing, some models to become unavailable in some regions, and provenance of training data to become a procurement question you are asked about by your own customers. The engineering response is to avoid deep coupling to a single model, which is good practice regardless.
  • If the output-similarity claims gain traction: the controls in the page on output ownership become more important — similarity checking before publication, human review on public-facing assets, and a clear-eyed reading of your provider’s indemnity.
  • If the defences hold broadly: very little changes operationally, and the practices above cost you almost nothing to have adopted.

That asymmetry is the actionable conclusion. The measures worth taking now are the ones that are cheap and useful in every branch, rather than a bet on an outcome.

Tracking it without a subscription

If you need to follow this, follow primary sources rather than commentary: court dockets and published judgments for the cases you care about, official journals and legislature sites for statutory changes, and the guidance pages of the relevant copyright offices, which are updated more often than people expect. Set a calendar reminder rather than relying on a feed, and when you read a summary — including this page — check what it is a summary of and when it was written.

Training Data Copyright: The Shape of the Dispute · Multigrid