Skip to content

The US Copyright Office's AI Report, Part 3: Training Data

10 min read · updated August 11, 2026

On 9 May 2025 the United States Copyright Office released Part 3 of its AI report, on generative AI training, as a pre-publication version. It is the most detailed official American analysis of whether training on copyrighted works is fair use, it reaches no single answer, and its standing has been contested since the day after it appeared.

The document and its unusual status

The release was labelled a pre-publication version, with a note that a final version would follow and that no substantive change to the analysis or conclusions was expected. The following day, 10 May 2025, the Register of Copyrights, Shira Perlmutter, was removed from office; she brought suit challenging the removal, and that litigation proceeded separately from anything to do with the report’s content.

The consequence for a reader is specific and worth stating plainly. The report is not a regulation and was never binding; agency reports persuade rather than bind. But its persuasive weight depends partly on its institutional standing, and a pre-publication document whose author was removed the next day carries a different weight in argument than a settled agency position. Litigants on both sides have cited it. Check the Copyright Office’s AI page for whether a final version has since been issued, and in what terms, before relying on the pre-publication text.

This page is not legal advice. Fair use is decided case by case on the facts of a specific use, an agency report does not decide it, and nothing here tells you whether your own training or fine-tuning is lawful. Take advice on your own facts.

What the Office says training does

Before reaching fair use, the report works through what acts of reproduction occur. Assembling a training corpus involves making copies. So does the ingestion and processing pipeline, and so, transiently, do the operations during training itself. The Office treats these as implicating the reproduction right, which is why the entire question ends up resting on exceptions rather than on whether copying occurred.

The report also addresses whether model weights themselves can constitute copies. Its position is nuanced: weights are not ordinarily a reproduction of the training works, but where a model has memorised particular works to the point that they can be retrieved from it, the analysis may differ. This is one of the places where the American and English approaches have diverged in practice, since the English High Court reached a firm conclusion on a related question in the Getty proceedings — see what the UK court decided about model weights.

The fair-use analysis

The report works through the four statutory factors in 17 U.S.C. §107 and declines to give one answer, holding instead that outcomes range across a spectrum depending on the use.

  • Purpose and character. Transformativeness is treated as a spectrum rather than a switch. Training for research, analysis or purposes far removed from the works’ expressive value sits at the strongly transformative end. Training a model to generate content that substitutes for the kind of works it was trained on sits at the other, and the report treats the deployed model’s purpose as relevant rather than looking only at the act of training. Whether the works were lawfully acquired is treated as bearing on this factor.
  • Nature of the work. Conventional and rarely decisive: expressive works weigh against fair use more than factual ones, and training corpora typically contain both.
  • Amount and substantiality. Training generally uses entire works, which ordinarily weighs against fair use, but the report accepts that copying a whole work can be reasonable where necessary to the purpose and where the work is not disclosed in the output. Guardrails that prevent regurgitation are relevant here.
  • Market effects. The report treats this as the heaviest factor and identifies three routes to harm: lost sales where outputs substitute for the originals, harm to the market for licensing works for training, and dilution of the market for the kind of work in question.

The report’s overall conclusion is deliberately unsatisfying and probably correct: some uses are clearly fair, some are clearly not, and a large middle range depends on facts that differ case by case. It also expresses a preference for voluntary and collective licensing markets over a statutory or compulsory licence, and does not recommend legislation.

Market dilution, the contested idea

The most novel and most criticised element is market dilution: the idea that flooding a market with machine-generated works of a kind can harm the market for human works of that kind, even where no particular output substitutes for any particular work.

Critics argue this stretches the fourth factor beyond harm to the market for the copyrighted work itself, which is what the statute names, into harm from competition — and that competition from new works is not ordinarily cognisable copyright harm. Supporters argue the fourth factor has always encompassed harm to potential and derivative markets, and that a model trained on a body of work to produce more of that kind of work is a straightforwardly derivative market effect. This is genuinely unresolved. It would be settled by appellate rulings addressing the theory directly, and the district courts that have engaged with it have not spoken with one voice.

What courts have done since

Two district court decisions in June 2025 are the ones most often set against the report, and both illustrate that a court’s fair-use answer turns on the record in front of it.

In the Northern District of California, Judge William Alsup held in June 2025 that training on lawfully purchased books was fair use, describing the use as highly transformative, while holding separately that maintaining a library of pirated copies was not excused by the training purpose. The distinction between how the copies were obtained and what they were used for is the reusable part of that reasoning; the case later settled on the pirated-library exposure. Days later, also in the Northern District of California, Judge Vince Chhabria granted summary judgment to the defendant on fair use for the named plaintiffs before him, while saying at length that market dilution may be the most important consideration in these cases and that the plaintiffs had lost because they failed to develop a record on it — an outcome that is closer to the Office’s framing than the headline suggested.

Filings and opinions in these matters are available through CourtListener’s docket archive. Read them rather than summaries: the reasoning is fact-bound, and none of these decisions establishes a general rule about training. For a case that went the other way on very different facts, see the Thomson Reuters v Ross ruling.