When a software vendor adds an artificial intelligence (AI) provider to its subprocessor list, the first thing that happens is entirely ordinary: a new subprocessor has been engaged, and the General Data Protection Regulation (GDPR) gives you the same rights and obligations as for any other. Notice under Article 28(2), a window to object, an update to your record of processing activities, and a transfer question if the provider sits outside the European Economic Area.
What is not ordinary is the set of questions the answer depends on. An AI subprocessor differs from a hosting provider in four specific ways, and each one can change whether the arrangement is acceptable. This article covers the ordinary part briefly, then the four questions worth actually asking.
The short answer
- An AI subprocessor is a subprocessor. Article 28(2) notice, your right to object, and Article 28(4) liability all apply unchanged.
- Your register changes: recipients under Article 30(1)(d), and destinations under Article 30(1)(e) if the provider processes outside the EEA.
- Four questions are specific to model processing: whether your data trains the model, how long it is retained and why, where inference actually happens, and whether outputs drive decisions about people.
- The first of those is the one that can change the legal analysis entirely, because a provider processing your data for its own purposes is not acting solely as a processor for that processing.
- The EU AI Act is a separate regime. A vendor's position under it does not answer an Article 28 question.
It is a subprocessor change first
Nothing about AI removes the ordinary machinery, and starting there prevents the analysis from becoming exotic.
Notice and objection. Under Article 28(2), a processor may not engage another processor without your prior specific or general written authorisation, and where the authorisation is general it must inform you of intended changes so you have the opportunity to object. Whether you get meaningful notice depends on which model your contract uses, covered in general vs specific authorisation for subprocessors, and what to do with the notice is covered in how to handle a subprocessor objection.
Equivalent obligations and liability. Article 28(4) requires your processor to bind the subprocessor to the same data protection obligations, and keeps your processor fully liable to you for the subprocessor's performance. That does not change because the subprocessor runs models.
Chain depth. AI providers have their own subprocessors, typically cloud infrastructure. Adding one usually adds a layer rather than a single party. The European Data Protection Board's Opinion 22/2024 on reliance on processors and sub-processors addresses this directly: to meet your Article 28 obligations you need available information on the identity of all processors and sub-processors in the chain, not only the party you contracted with.
Register and transfers. The new party is a recipient for Article 30(1)(d) purposes, and if it processes outside the EEA the destination belongs in Article 30(1)(e). Which of your records that touches is covered in what Article 30 actually requires, and whether a new destination pulls a transfer assessment into scope is covered in when a Transfer Impact Assessment is required.
Question 1: is your data used to train or improve the model?
This is the question that can move the arrangement outside the processor relationship altogether.
A processor processes only on the controller's documented instructions. If the AI provider uses the data you send it to train or improve its own models, it is processing for its own purposes, and for that processing it is not acting as your processor. The consequences are not cosmetic: there would need to be a lawful basis for that separate processing, transparency obligations toward the individuals concerned, and a role allocation that your existing Article 28 contract does not describe.
In practice, enterprise and API tiers of major providers commonly commit not to train on customer content, while consumer tiers frequently reserve the right to. The distinction between tiers of the same product is therefore material, and "we use provider X" is not an answer. What you need is which tier, under which terms, and whether training is excluded by contract rather than by policy statement.
The EDPB's Opinion 28/2024 on AI models is the reference point for the wider questions this raises, including when a model can be considered anonymous and how legitimate interest may be assessed as a basis for developing and deploying models. Its relevance to a vendor review is indirect but useful: it establishes that model training on personal data is a processing operation requiring its own analysis, not an incidental technical detail.
Question 2: how long is your data retained, and why?
Model providers frequently retain inputs and outputs for a period for abuse monitoring and safety review, even where they do not train on them. That retention is a distinct purpose from delivering the service, and it is often operated by human reviewers rather than automatically.
Two conflicts arise. The retention window may exceed what your own register documents for that data, which makes your Article 30 retention field inaccurate. And human review introduces access by people at a party two steps removed from you, which is a confidentiality question your Article 32 assessment should account for.
Ask for the retention period, the purpose, whether human review occurs, and whether a zero-retention or reduced-retention configuration is available for your tier. Providers increasingly offer one, and it is usually a setting rather than a negotiation.
Question 3: where does inference actually happen?
A vendor's stated hosting region describes where it stores your data. It does not necessarily describe where a model call is processed. Inference may be served from a different region, and it may fail over to another one under load.
Because remote access from a third country is itself a transfer, this is not answered by knowing where data is stored. What you need is the processing region for model calls, whether it is contractually pinned or best-effort, and what happens on failover. A vendor that can name the region and commit to it contractually has answered the question; a vendor that describes its storage region has not.
Where the answer places processing outside the EEA, the ordinary Chapter V analysis applies: which transfer tool covers the leg, and whether an assessment is required behind it.
Question 4: does the output drive a decision about a person?
Most AI subprocessors added by software vendors do unremarkable things: summarising a support ticket, drafting a reply, classifying a document. Some do not.
Where model output feeds a decision about an individual, two further provisions come into view. Article 22 restricts decisions based solely on automated processing that produce legal effects or similarly significantly affect the person. And Article 35 makes a data protection impact assessment likely where processing involves systematic and extensive evaluation of personal aspects, or the innovative use of new technology.
The practical test is whether a human meaningfully reviews the output before it affects anyone, and whether that review is real or a rubber stamp. Where the answer is that the output is advisory and a person decides, the analysis stays simple. Where the output effectively is the decision, the review needs to be considerably deeper than a subprocessor notice review.
The AI Act is a separate regime
The EU AI Act imposes its own obligations, on its own timeline, allocated by role in the AI value chain. It sits alongside the GDPR rather than replacing it.
The practical consequence for a vendor review is narrow but worth stating: a vendor's AI Act posture does not answer a GDPR question. Conformity work under one regime is not evidence of a lawful basis, a valid transfer tool, or an adequate Article 28 contract under the other. Keep the two assessments separate, and do not accept documentation from one as closing a gap in the other.
How to review one in practice
When a subprocessor notice names an AI provider, a proportionate review answers eight things:
- Which product and tier of the provider is in use.
- Whether training or model improvement on your data is contractually excluded.
- The retention period for inputs and outputs, and its purpose.
- Whether human review of content occurs, and by whom.
- Whether a reduced-retention or zero-retention configuration is available and enabled.
- The processing region for inference, and whether it is contractually pinned.
- Which transfer tool covers the leg if processing occurs outside the EEA.
- Whether output feeds any decision about an individual, and what human review sits in front of it.
Record the answers with the date and the source, because these terms change more often than infrastructure terms do. An answer from six months ago about a provider's retention default is not a current answer.
Common mistakes
- Treating it as a technology question rather than a subprocessor question. The Article 28 machinery applies first, and skipping it means skipping your objection right.
- Accepting the provider's name as the answer. Terms differ substantially between tiers of the same product, and the tier determines whether training is excluded.
- Assuming the storage region covers inference. They are separate facts and vendors often state only the first.
- Confusing a policy statement with a contractual commitment. A published position on training can change; a contract term is enforceable.
- Forgetting the register. A new subprocessor and possibly a new destination country both belong in the Article 30 record, and neither updates itself.
- Stopping at the AI provider. Its own subprocessors are part of the chain you are accountable for understanding.
Frequently asked questions
Does adding an AI subprocessor require a new DPIA?
Not by itself. A DPIA is required under Article 35 where processing is likely to result in a high risk, and adding a model provider that summarises text does not usually change the risk profile. It becomes likely where the processing involves systematic evaluation of personal aspects, decisions affecting individuals, large-scale special-category data, or genuinely novel use of technology. Where a DPIA already exists for the activity, Article 35(11) requires review if the risk has changed.
Can we object to an AI subprocessor specifically?
Where your contract uses general authorisation, Article 28(2) gives you the opportunity to object to an intended change, and that right is not limited by the type of subprocessor. What follows from an objection depends on the contract: some provide for termination of the affected service, others leave you negotiating. The practical constraint is the notice period, which is why it is worth fixing before you need it.
Is a model provider a processor or a controller?
It depends on the processing, not on the company. For providing inference on your instructions under terms that exclude other uses, it acts as a processor in the chain. For training its own models on the same data, it would be pursuing its own purpose, which is not processor activity. A single provider can occupy both positions across different processing, which is why the training question has to be answered explicitly.
Does anonymised or pseudonymised input remove the issue?
Genuine anonymisation takes the data outside the GDPR, but the threshold is high and it is assessed against realistic re-identification, not intent. Pseudonymisation does not: pseudonymised data remains personal data, and it reduces risk rather than removing the obligation. Free-text content sent to a model is particularly resistant to reliable anonymisation, because identifiers appear in prose rather than in fields.
How do we keep up with these changes?
Subprocessor lists and AI terms change on the vendor's schedule rather than yours, and the changes are frequently the kind that a general authorisation notice mentions in a line. Watching the published subprocessor and legal pages on a recurring basis, and keeping dated evidence of what they said, is the only mechanism that does not depend on someone remembering to check.
Where DPAFlow fits in
DPAFlow does not evaluate AI providers and does not provide legal advice. What it does is make sure the change reaches you, and that you can prove what a vendor published and when.
It monitors vendor subprocessor lists, Data Processing Agreement pages, and related legal sources on a recurring schedule, captures dated evidence of what each source said, and reports what changed between captures rather than only that something did. Detected entries carry the evidence they were extracted from and are reviewed before they reach a customer workspace. The product overview covers how that capture and evidence chain works.
Because a detected change on a vendor a record depends on flags the affected processing activities and transfer assessments for re-review, with the triggering event recorded against them, the eight review questions above get asked when there is something new to ask them about. What the answers mean for your organisation stays a judgement for your team.