Skip to content

The EU AI Act's Ban on Scraping Faces for Recognition Databases

8 min read · updated August 11, 2026

Article 5(1)(e) has no exceptions, no law enforcement carve-out and no conditions. It is also narrower than it first appears, because every limiting word in it — untargeted, facial images, internet or CCTV, create or expand — is load-bearing.

The shortest prohibition in the Act

Article 5(1)(e) of Regulation (EU) 2024/1689 prohibits the placing on the market, putting into service for this specific purpose, or use of AI systems that create or expand facial recognition databases through the untargeted scraping of facial images from the internet or CCTV footage.

Compare it with its neighbours. Article 5(1)(f), on emotion recognition, carries a medical and safety exception. Article 5(1)(g), on biometric categorisation, carves out lawfully acquired datasets. Article 5(1)(h) has three enumerated law enforcement objectives and an authorisation regime running to several further paragraphs. Point (e) has nothing. If conduct is inside it, it is prohibited, including for national police forces acting under national law.

One structural caveat: Article 2 excludes AI systems placed on the market, put into service or used exclusively for military, defence or national security purposes from the scope of the whole Regulation. That exclusion operates before Article 5 is reached, so it is not an exception to point (e) so much as a limit on the instrument itself, and its own boundaries are contested.

Not legal advice. Article 5 breaches carry the highest penalty tier under Article 99(3) — up to EUR 35,000,000 or 7% of total worldwide annual turnover — and a facial dataset is also personal data processing, and biometric data processing, under the GDPR. If you are assessing a real dataset or pipeline, take advice on the facts.

What “untargeted” excludes

The prohibition is on untargeted scraping, and that word does the work of every missing exception. Collection directed at identified individuals, on a defined and lawful basis, is not untargeted: an investigator lawfully gathering images of a named suspect is doing something the provision does not describe.

The mischief is indiscriminate accumulation — sweeping in whoever happens to appear, with no criterion tying the collection to any particular person or purpose at the moment of collection. Scale is the usual signal but it is not the test. A small indiscriminate sweep is untargeted; a large but individually justified collection is not.

There is no threshold in the text and no guidance that supplies one. Where the line falls for an intermediate case — a criterion that identifies a class rather than individuals, say — is unresolved, and the honest answer is that it will be worked out through national enforcement.

Two sources, and only two

The provision names the internet or CCTV footage. That is a closed list, and the omissions are consequential.

  • A licensed stock photo archive is not the internet or CCTV. Building a face database from a commercially acquired image corpus is not within point (e) on the plain words, whatever it is under the GDPR or copyright law.
  • Images uploaded by users to your own service are neither. The provision addresses harvesting from external sources, not accumulation through a service the subjects used.
  • Body-worn camera and vehicle camera footage are not obviously CCTV. Closed-circuit television has a settled ordinary meaning centred on fixed surveillance installations. Whether the term stretches to mobile capture is untested, and a functional reading that includes it is at least arguable given the recital’s mass surveillance rationale.

Note also “facial images” specifically. A database of gait, voice or iris data assembled by untargeted scraping is not caught by this provision. That looks like an oversight rather than a decision, but the text is the text.

“Create or expand” catches incremental growth as well as initial construction, so an existing database maintained by continued scraping remains in breach for as long as the scraping continues. That is what makes the prohibition bite on ongoing operations rather than only on historic acts.

Where the provision came from

Recital 43 sets out the rationale in terms: such practices add to the feeling of mass surveillance and can lead to gross violations of fundamental rights, including the right to privacy. The recital does not name a company, but the practice it describes — assembling a face-search database of billions of images scraped from the public web — is a specific one, and it was the subject of enforcement action by several European data protection authorities before the AI Act was adopted. Those decisions were taken under the GDPR, and their enforcement history is the reason a separate prohibition was thought necessary: the substantive breach was found repeatedly, and the practice continued.

That history is documented at the running tally of European facial recognition enforcement. The point for reading Article 5(1)(e) is that the drafters were legislating against a known, ongoing practice rather than a hypothetical, which is why the provision is short and admits no balancing.

What the GDPR does that this does not

Article 5(1)(e) prohibits building the database. It says nothing about using one, about the model trained on it, or about the outputs. Those questions are answered elsewhere.

Facial images processed for the purpose of uniquely identifying a natural person are biometric data and a special category under GDPR Article 9(1), which prohibits processing unless one of the Article 9(2) conditions applies. Consent and substantial public interest are the candidates, and for indiscriminate web scraping neither is straightforward. The GDPR therefore already bars much of this conduct for anyone within its territorial scope, and it reaches the processing rather than only the construction.

On the AI Act side, remote biometric identification systems are separately classified as high-risk under Annex III point 1(a), so a system that identifies people from a lawfully assembled database carries the Chapter III obligations. Real-time use in publicly accessible spaces for law enforcement is separately restricted by Article 5(1)(h) and its narrow exceptions, and the Annex III biometrics category sets out what the permitted systems owe.

The layering is deliberate. One instrument prohibits the collection method, another regulates the identification system, a third governs the personal data throughout. A compliance answer that stops at Article 5(1)(e) has cleared one of three.