Web-Scraped Training Data and GDPR's Lawful Basis Question
10 min read · updated August 11, 2026
A crawl of the public web collects personal data about millions of people who will never be told about it. That is not automatically unlawful under the GDPR, and it is not automatically lawful either. The whole question runs through one provision and one balancing test.
Why it is legitimate interest or nothing
Article 6(1) offers six bases. Five of them are unavailable to a general web crawl on inspection. Consent is impossible at the scale and would have to be obtained before collection from people whose identity you learn only by collecting. Contract fails because there is no contract with the people in the corpus. Legal obligation, vital interests and public task do not describe commercial model training. What remains is Article 6(1)(f), the legitimate interests of the controller or a third party, except where overridden by the interests or fundamental rights and freedoms of the data subject.
That is not a weak basis. The Court of Justice confirmed in Case C-621/22, decided on 4 October 2024, that a purely commercial interest can be a legitimate interest within the meaning of Article 6(1)(f); the question is always the balancing, not the legitimacy of wanting to make money. Case C-252/21, Meta Platforms v Bundeskartellamt, decided on 4 July 2023, is the other essential reference, because it deals with the reasonable expectations of a person whose data is used for purposes far removed from the context of collection.
The three limbs, applied to a crawl
The test has three limbs and they are assessed in order. The purpose limb asks whether the interest is legitimate, specific and real rather than speculative. “Training AI” fails as a statement of interest because it names an activity rather than an interest; “developing a conversational assistant for X, which requires a corpus with property Y” is the level of specificity the assessment is looking for. A stated interest that is really a placeholder for “everything we might build later” also collides with purpose limitation under Article 5(1)(b).
The necessity limb asks whether the processing is necessary for that interest and whether a less intrusive means would achieve it. This is where a scraper is most exposed, because the honest answer is often that a smaller, filtered or licensed corpus would have worked less well rather than not at all. Necessity does not mean indispensable, but it does mean you have to have considered the alternatives, and the consideration has to be recorded. Filtering out categories of site before ingestion is the single cheapest way to strengthen this limb.
The balancing limb asks whether the data subject’s interests or fundamental rights override. For a crawl the aggravating factors cluster: the data subject has no relationship with the controller, no notice, no realistic expectation of the use, and no practical way to object before the fact. The volume is enormous. The corpus will contain special-category data under Article 9 by accident — forum posts about illness, political affiliation, sexual orientation — for which Article 6 is not enough on its own, and for which no Article 9(2) condition obviously fits public web content, since the “manifestly made public by the data subject” condition in Article 9(2)(e) requires a deliberate act by that person and cannot be assumed from mere accessibility.
What EDPB Opinion 28/2024 added
The European Data Protection Board adopted Opinion 28/2024, on certain data protection aspects related to the processing of personal data in the context of AI models, on 17 December 2024, following an Article 64(2) request from the Irish supervisory authority. Three things in it change how this assessment is written.
First, on anonymity: the Board declined to accept that a trained model is anonymous merely because it does not store records. Anonymity has to be demonstrated, with reference to the likelihood of extracting personal data from the model and of obtaining it through queries, and the assessment is case by case. That matters because if the model is not anonymous, the GDPR follows the model rather than stopping at the training set.
Second, on legitimate interest: the Board set out how the three-limb test applies to development and to deployment as separate processing operations, and gave weight to whether the data subject could reasonably expect the use, to the source and nature of the data, and to mitigating measures.
Third, and most consequential, on unlawful origin: the Board addressed what happens where a model is developed on unlawfully processed personal data and then deployed by the same or another controller, and treated the lawfulness of the earlier stage as capable of affecting the later one. A downstream deployer cannot always treat the provenance of a model as somebody else’s problem. The practical effect is that provenance questions belong in vendor due diligence, which is the subject of the subprocessor checklist.
The Garante’s dated decisions
Italy’s Garante per la protezione dei dati personali has been the most active regulator on this specific ground. In May 2024 it published a guidance document on artificial intelligence and web scraping, addressed to operators of websites that publish personal data, setting out measures they may take against scraping — dedicated areas behind registration, rate limiting, bot-management measures, robots.txt and terms-of-service signals. It is worth reading for what it implies about the scraper’s position as much as for its stated audience: it treats the publisher’s expressed refusal as meaningful, which feeds directly into the reasonable-expectations part of the balancing.
In December 2024 the same authority concluded its proceeding against OpenAI arising from the ChatGPT suspension of March 2023, imposing a fine of €15 million together with a corrective measure requiring a public information campaign, and finding breaches including the absence of an adequate legal basis and a failure of transparency towards data subjects. OpenAI stated it would appeal. The primary record for both is the Garante’s own decision register; the amount as announced is not necessarily the amount that survives appeal, and the dated position is discussed further in the enforcement round-up.
The levers that move the balance
- Respect expressed refusals at crawl time. A crawler that honours robots.txt and machine-readable opt-outs is in a materially different position on reasonable expectations from one that does not, quite apart from the separate copyright question under the Article 4 CDSM text-and-data-mining opt-out.
- Exclude source categories before ingestion. Sites whose subject matter is inherently special-category — health support forums, dating, political membership — are cheaper to exclude by domain than to filter by content afterwards, and the exclusion list is evidence.
- Filter and pseudonymise inside the pipeline. Direct identifier stripping, deduplication of memorisable passages and output filters are all mitigations the balancing test can weigh, and they interact with the anonymisation question.
- Publish the Article 14 notice anyway. Article 14(5)(b) can excuse individual notification where it would involve disproportionate effort, but it requires appropriate measures instead, including making the information publicly available. A controller that relies on the exemption and then publishes nothing has taken the exemption without the condition attached to it.
- Make the objection route real. Article 21(1) gives an absolute-in-form right to object to Article 6(1)(f) processing, which the controller may resist only on compelling legitimate grounds. What you can actually do once a model is trained is the uncomfortable subject of rights requests against a fine-tuned model.
None of this makes the question settled. Whether a general-purpose crawl for commercial model training passes the balancing test at all has not been decided by the Court of Justice, and national authorities have taken visibly different tones about it. Write the assessment as though it will be read by a regulator who starts from the view that it does not, and you will at least be arguing on the record you built rather than on the one you wish you had.