The Questions a DPIA Requires You to Ask an AI Vendor
10 min read · updated August 11, 2026
A DPIA is your document and your legal obligation, but half the facts it needs are the vendor’s. This is the list, built so that every question names the part of Article 35(7) it exists to fill — which is what turns an unanswered question into a visible gap rather than a blank line.
Why the vendor is on the critical path
Article 35(1) GDPR requires a data protection impact assessment where processing, in particular using new technologies, is likely to result in a high risk to the rights and freedoms of natural persons. Article 35(3) lists three cases where one is required in particular, and the first — systematic and extensive evaluation of personal aspects based on automated processing, including profiling, on which decisions producing legal or similarly significant effects are based — is the one most AI deployments in decision-making contexts sit near. Primary text: Regulation (EU) 2016/679 on EUR-Lex. The triggers page covers when you are in and out.
The obligation is the controller’s and cannot be delegated. But Article 35(7) requires you to describe the processing, assess necessity and proportionality, assess risk and describe mitigations — and for a hosted model you do not know the retention, the sub-processors, the log handling or the human review path without asking. The AI Act makes the same connection from the other side: Article 26(9) directs a deployer of a high-risk system to use the Article 13 information to carry out its DPIA.
There is a threshold question worth settling before you send anything. The Article 29 Working Party’s DPIA guidelines, WP248 rev.01, as endorsed by the EDPB, set out nine criteria — evaluation or scoring, automated decision-making with legal or similar effect, systematic monitoring, sensitive data, large scale, matched or combined datasets, vulnerable data subjects, innovative use of technology, and processing that prevents data subjects from exercising a right or using a service — and treat two or more as generally indicating a DPIA is required. Most generative deployments involving personal data hit “innovative use” on its own, so the practical question is which second criterion applies. Note that hitting only one does not mean no DPIA is needed; it means you should record why you concluded one was not.
The four limbs, and what each needs from them
- 35(7)(a) — systematic description of the processing and its purposes. Needs: what personal data leaves your boundary, where it goes, who touches it, how long it stays, and what else it is used for. Vendor-dependent almost entirely.
- 35(7)(b) — necessity and proportionality. Needs: whether a less intrusive configuration exists — regional processing, zero retention, no human review, redaction before transmission. You cannot argue proportionality if you did not know the vendor offered a narrower option.
- 35(7)(c) — risks to rights and freedoms. Needs: accuracy characteristics, failure modes, memorisation and regurgitation risk, the effect of the vendor’s abuse-monitoring on confidentiality, and the realistic exposure of a breach.
- 35(7)(d) — measures to address the risks. Needs: the actual controls, the contractual commitments behind them, and the evidence that they operate.
The question set
Send these as written, with the article reference attached to each. Vendors answer a referenced question more precisely than an open one, and the reference is what makes a deflection legible.
A. Data flow and retention [35(7)(a)]
A1 What is retained from a request: prompts, outputs,
attachments, embeddings, metadata? State each separately.
A2 Retention period for each, in days, by default and by
configuration. Is zero retention available, and on which plan?
A3 Is content retained for abuse or safety monitoring even where
retention is otherwise off? For how long, and who can read it?
A4 Can a human employee or contractor read request content?
Under what trigger, with what approval, with what logging?
A5 Which geographic regions process and store content? Is region
pinning contractual or best-effort?
A6 Is customer content used to train, fine-tune, evaluate or
improve any model? Confirm as a contractual commitment, not a
policy page.
B. Parties [35(7)(a), Art 28]
B1 Current sub-processor list: entity, function, country, and
whether it processes content or only metadata.
B2 For which processing, if any, do you act as controller rather
than processor? Name the purposes.
B3 Notice period for sub-processor changes, and the objection
mechanism.
B4 Transfer mechanism per recipient outside the EEA/UK, and the
date of the most recent transfer impact assessment.
C. Alternatives and minimisation [35(7)(b)]
C1 What configuration minimises personal data processed —
redaction, field exclusion, ephemeral mode, on-region only?
C2 Which of these degrade the service, and how?
C3 Is there a deployment model where content does not leave our
infrastructure or our cloud tenancy?
D. Risk characteristics [35(7)(c)]
D1 Documented accuracy metrics and the population measured on.
D2 Known failure modes affecting individuals: fabrication of
personal details, group performance disparities, prompt
injection leading to disclosure.
D3 Can content submitted by one customer appear in another
customer's output? Explain the mechanism that prevents it.
D4 For fine-tuning: what would it take to remove one person's
data from a trained artefact, and what is your answer to an
erasure request that reaches it?
E. Controls and evidence [35(7)(d)]
E1 Certifications held, with scope statements and dates —
ISO 27001 SoA, ISO 42001, SOC 2 Type II report.
E2 Breach notification commitment to us, in hours, and the
content of the notice.
E3 Support for data subject rights: access, rectification,
erasure, restriction — what can you execute, in what time?
E4 Audit rights: scope, frequency, trigger, and whether findings
may be shared with our supervisory authority.
E5 Model version pinning, deprecation notice period, and
notification of behaviour-affecting changes.Reading the answers
Four patterns are worth recognising because they look like answers and are not.
The policy-page answer. A link to a trust centre in response to a contractual question. Trust pages change without notice and bind nobody. The follow-up is one sentence: “please confirm this as a term of the agreement or a signed side letter”.
The abuse-monitoring gap. A vendor will often answer A2 with zero retention and A3 with a thirty-day safety retention, and the two answers appear on different pages. The second is the one your DPIA has to describe, because it is a real retention of real content with real human access. Ask A3 separately for exactly this reason.
The controller shrug. B2 is the question most often answered vaguely, and it matters: if the vendor is an independent controller for some processing, your Article 28 analysis does not cover it, your DPIA has to describe a disclosure to a third party rather than a processing on your behalf, and the lawful basis question reopens.
The erasure non-answer. D4 has no comfortable answer for a fine-tuned model and honest vendors say so. An answer that promises full erasure from a trained artefact without explaining the mechanism is more concerning than one that says the model is retrained on a schedule and the interim position is documented — see data subject requests against a fine-tuned model.
The certification cover page. E1 asks for scope statements deliberately. An ISO 27001 certificate names a scope, and a scope covering a corporate function rather than the production inference estate tells you nothing about the service you are buying. The same applies to a SOC 2 Type II: the value is in the exceptions section and the period covered, not in the fact that a report exists. A vendor that sends the certificate rather than the scope statement is answering a different question, and the polite follow-up is to ask for the statement of applicability or the scope paragraph by name.
After the answers come back
Write the description first, from the answers, before you assess anything — if you cannot write a paragraph describing where the data goes, you do not yet have enough to assess. Then run necessity and proportionality against the alternatives in section C, and record the alternatives you rejected and why; that record is the substance of the 35(7)(b) limb and the part most DPIAs skip.
Where residual high risk remains after mitigation, Article 36 requires prior consultation with the supervisory authority before processing. That is a real gate with a real timeline, and finding it late is how a launch date slips. Consider seeking the views of data subjects or their representatives where appropriate under Article 35(9) — in an employment context, this is where works council engagement belongs. Structure for the generative case specifically is in the generative AI DPIA structure page, and the separate question of whether you may appoint the vendor at all is in the Article 28 sub-processor page.