Skip to content

Real-World Testing of High-Risk AI Outside a Sandbox (Article 60)

10 min read · updated August 11, 2026

Testing a high-risk system on real subjects before it is placed on the market is permitted, and it is permitted outside a regulatory sandbox. Article 60 sets the price: an approved plan, a registration, informed consent, a hard time limit, and the ability to reverse anything the system decides.

Which systems Article 60 is available for

Article 60(1) of Regulation (EU) 2024/1689 opens the route to providers and prospective providers of high-risk AI systems listed in Annex III, in accordance with the article and with a real-world testing plan, and without prejudice to the prohibitions in Article 5. Three limits are packed into that sentence.

  • Annex III only. Systems that are high-risk because they are safety components of products covered by the Union harmonisation legislation in Annex I are not within Article 60. Those products have their own sectoral testing regimes — clinical investigations under the medical devices framework, type-approval testing for vehicles — and the AI Act does not add a parallel one for them.
  • Before placing on the market. The article governs testing prior to placing on the market or putting into service. Once a system is on the market, you are in post-market monitoring under Article 72, not real-world testing.
  • Article 5 still bites. Nothing in Article 60 authorises testing a practice prohibited by Article 5. There is no experimental exception to the prohibitions.

Chapter VI, which contains Articles 57 to 63, applies from 2 August 2026 under Article 113. The provisions are published at EUR-Lex.

Real-world testing of an AI system on people frequently engages other law at the same time — data protection, medical research governance, employment law, sectoral rules. Article 60 compliance is not a substitute for any of it, and this page is not legal advice. Take advice on your own facts before testing on real subjects.

The sequence, and the gate at each step

Article 60(4) lists the conditions cumulatively, but in practice they form an order, and each step is a gate that stops the next one.

  1. Draw up a real-world testing plan and submit it to the market surveillance authority in the member state where the testing is to be conducted. The Commission is to adopt implementing acts specifying the detailed elements of the plan, so its required contents are set centrally rather than by each authority.
  2. Obtain the authority’s approval of the testing and the plan. Where the authority has not provided an answer within 30 days, the testing and the plan are deemed approved — a tacit approval mechanism, subject to national law providing otherwise. Plan for the 30 days, not for the possibility of silence.
  3. Register the testing in the EU database under Article 71(4), with a Union-wide unique single identification number and the information set out in Annex IX. For testing of systems in the biometrics, law enforcement and migration areas of Annex III, that registration is in the non-public section of the database. The identification number is not administrative trivia: it has to be given to every subject as part of the consent information.
  4. Be established in the Union or appoint a legal representative established in the Union. A third-country provider cannot run an Article 60 test without a Union-side counterparty, which parallels the authorised representative requirement for placing systems on the market.
  5. Obtain informed consent from the subjects under Article 61, before they participate. The narrow exception is discussed below.

Two further conditions in Article 60(4) constrain the shape of the test rather than the paperwork. The testing may not last longer than necessary and in any event no more than six months, extendable once by a further six months where the provider notifies the market surveillance authority in advance with a justification. And subjects who are vulnerable persons due to their age or disability must be appropriately protected. Where the provider organises the testing with one or more deployers, those deployers must be informed of all aspects relevant to their decision to participate and given the relevant instructions for use, and the roles and responsibilities must be allocated in an agreement.

Article 61 sets out what informed consent means here, and it is deliberately close in structure to research ethics practice rather than to a data protection consent banner. It must be freely given prior to participation, after the subject has been duly informed with concise, clear, relevant and understandable information about the nature and objectives of the testing and the possible inconvenience; the conditions under which the testing is conducted, including the expected duration of participation; the subject’s rights and guarantees, including the right to refuse and to withdraw at any time without any resulting detriment and without needing to provide a justification; the arrangements for requesting the reversal of the predictions, recommendations or decisions of the system, or for having them disregarded; and the Union-wide unique single identification number of the testing together with the contact details of the provider or its legal representative. Consent must be dated and documented, and a copy given to the subject or their legal representative.

The right to withdraw operates without detriment, without justification, at any time, and carries a right to request immediate and permanent deletion of the subject’s personal data. Withdrawal does not affect activities already carried out. That last clause is worth designing around: a system whose training corpus is updated from test interactions has to be able to answer what “already carried out” means for data that has been used to change weights, and that is the same hard question as consent withdrawal against a trained model generally.

There is one exception, and it is narrow. Article 60(4) allows testing in the law enforcement context to proceed without informed consent where seeking it would prevent the system from being tested, on condition that the testing and the outcome of the testing do not have any negative effect on the subjects and their personal data is deleted after the test is performed. Both conditions are cumulative and the first is strict: not “minimal effect”, not “proportionate effect”, no negative effect.

Duties while the test is running

Article 60(4) requires the testing to be effectively overseen by the provider or prospective provider and by persons with appropriate qualifications, training and the necessary resources, and requires that the predictions, recommendations or decisions of the system can be effectively reversed and disregarded. That reversibility condition is the substantive heart of the article. It means a real-world test cannot be run on a system whose outputs take irreversible effect — if a decision cannot be unwound, the test is not lawful under this route however good the consent process is.

If a serious incident occurs, Article 60 requires immediate suspension of the testing until mitigation takes place, and failing that, termination; the provider must have an effective procedure for the immediate recall of the system. Serious incidents identified in the course of the testing are reported to the national market surveillance authority under the Article 73 reporting regime. The provider also notifies the authority of any suspension or termination and of the final outcomes.

The boundary with ordinary product testing

The question this page gets asked most is where ordinary internal testing ends and Article 60 begins, and the Regulation does not draw a bright line. What it gives you is the definition of testing in real-world conditions: temporary testing of a system for its intended purpose in real-world conditions outside a laboratory or otherwise simulated environment, with a view to gathering reliable and robust data and assessing and verifying conformity. Testing on synthetic data, on historical records, or on staff acting as users in a controlled setting is not that. Running the system against live cases, affecting real people, for its intended purpose, is.

The uncomfortable middle is the shadow deployment: the system runs on real cases in parallel with the human process, and its outputs are recorded but not acted on. It is genuinely arguable whether that is testing in real-world conditions within the definition, since no prediction affects any subject. The counter-argument is that subjects are real, their data is processed for the intended purpose, and the purpose is verifying conformity — which is the definition. It is not settled, and the conservative reading is that a shadow deployment on identifiable individuals is safer treated as Article 60 testing than argued out of it after the fact. What would settle it is guidance from the Commission or a market surveillance authority taking a position; neither has produced one.