Why Art. 6 GDPR also applies to AI training
Anyone training an AI model on data relating to identified or identifiable individuals is processing personal data within the meaning of the GDPR – regardless of whether the training serves an internal analytics tool, a chatbot or a recommendation system. Art. 6(1) GDPR requires at least one of the listed legal bases for any such processing. Without one, the training is unlawful – even if the trained model itself no longer outputs personal data afterwards. The EU AI Act does not change this: its deadlines (such as the transparency obligations under Art. 50, applicable from 02.08.2026) govern other aspects of AI deployment and do not replace the data protection assessment required under Art. 6 GDPR.
Which legal bases actually hold up in practice
Art. 6(1) GDPR sets out six possible legal bases. In practice, the following are most relevant for AI training:
- Consent (lit. a): Requires consent “for one or more specific purposes”. Blanket clauses intended to cover “any future use for AI purposes” generally fail to meet this specificity requirement.
- Performance of a contract (lit. b): Only applies where the training is genuinely necessary to perform the contract concluded with the data subject – not where training data merely arise as a by-product of a contractual relationship.
- Legitimate interest (lit. f): The basis most commonly relied on in practice for training on existing customer or usage data. It requires a balancing test: the interests of the controller (or a third party) must not be overridden by the interests or fundamental rights of the data subject – a stricter standard applies where children are affected.
A legal obligation (lit. c) or a task carried out in the public interest (lit. e) is only available where a corresponding legal basis exists under Union or Member State law that sufficiently specifies the purpose, categories of data and retention period (Art. 6(3) GDPR).
The typical gap: change of purpose for existing data
The most common error in practice is not the absence of a legal basis for collection, but rather for further processing: data was originally collected for a different purpose – for example, contract performance or customer support – and is later used to train an AI model. Art. 6(4) GDPR requires a compatibility assessment in such cases, unless fresh consent or a statutory basis is available. This assessment must consider, among other things:
- the link between the original purpose of collection and the training purpose,
- the context in which the data was collected, in particular the relationship between the data subject and the controller,
- the nature of the data (for example, special categories under Art. 9 GDPR),
- the possible consequences of the further processing for the data subjects,
- existing safeguards such as pseudonymisation or encryption.
Where this assessment is not documented, there is no evidence, should it be needed, that the further processing was permissible – irrespective of how carefully the model was trained from a technical standpoint.
Web scraping and third-party data: particularly prone to errors
When training on externally sourced or scraped datasets, a solid legal basis is frequently missing altogether: the data subjects have not consented, there is no contractual relationship, and a balancing test under lit. f was never carried out or documented. Particularly with large-scale scraping from publicly accessible sources, it is often overlooked that “publicly accessible” is not the same as “lawful to process further”. Every data source feeding into a training set requires its own assessment under Art. 6 GDPR – blanket assumptions of lawfulness will not hold up in a dispute.
What this means for your organisation
Anyone training AI models on personal data should document, for each data source and each training run, which legal basis under Art. 6 GDPR applies and – in the event of a change of purpose – why the further processing is compatible under Art. 6(4) GDPR. This documentation is not an end in itself; it is the foundation for being able to account for your processing to supervisory authorities or data subjects if the need arises.
If you are unsure how your organisation currently stands with regard to AI training and data protection: the free risk check at /einstufung gives you an initial assessment of your obligations.