Click to open contact form.
Your Global Partners in the Business of Innovation

EDPB Adopts Landmark Guidelines on Anonymization and Web Scraping for Generative AI

General / July 30, 2026

Written by: Haim RaviaDotan Hammer

On July 8, 2026, the European Data Protection Board (EDPB) — the EU’s panel of data protection regulators — adopted two draft guidance documents of major significance and released them for public consultation until October 30, 2026. The first, Guidelines 02/2026 on anonymization, addresses the anonymization of personal data and is intended to replace the existing 2014 opinion on the subject, incorporating several important clarifications following rulings of the Court of Justice of the European Union (CJEU) and technological developments. The second, Guidelines 03/2026, provides new guidance on the scraping of personal data for the purpose of developing and training artificial intelligence systems.

On anonymization, the EDPB explains that the assessment of anonymity may differ from one entity to another: information considered anonymous for one entity may be identifiable for another. The guidance sets out two questions for determining whether information is anonymous — first, whether the information relates to a natural person, and second, whether that person is “identified or identifiable,” meaning whether they can be distinguished from others by using a direct identifier or through a combination of data points or unique identifiers — and emphasizes that identification should be examined from the controller’s perspective rather than the processor’s. It outlines two approaches that may be combined: a “contextual approach,” which examines the ability of each relevant party (including the controller, the data subject’s close circle, law enforcement and intelligence agencies, and even hackers) to re-identify the individual; and a “simplified approach,” which examines the theoretical ability of any party to do so. In assessing whether re-identification by reasonable means is possible, controllers must consider not only means accessible through third parties but also “chains” of means that several parties could combine, and the EDPB cautions against relying on contractual restrictions against re-identification or the absence of an apparent motive.

The guidance provides that data satisfying three cumulative criteria may be regarded as anonymous: no record isolation (meaning the data does not include a unique combination of values relating to an individual); no linkage (meaning the data cannot be linked to a record in another database relating to the same individual); and no inference (no unique and meaningful conclusion can be drawn from the data, where “unique” refers to an identified individual and “meaningful” implies an impact on that individual’s rights, rather than merely being derived from general or aggregate population data). Data may still qualify as anonymous where the risk of re-identification is negligible. The EDPB further clarifies that the anonymization process itself constitutes an act of processing, thereby requiring a lawful basis under Article 6 of the GDPR and appropriate documentation. Furthermore, anonymous data stored in a database that combines both anonymous and identifying data should be treated as personal data as long as the database cannot be properly separated.

Guidelines 03/2026 apply to private entities that scrape personal data from external sources for the purpose of developing and training generative AI (GenAI). The roles of the organizations involved must be assessed on a case-by-case basis. The party conducting the scraping is not necessarily the controller and may function as a processor where it acts on the AI developer’s documented instructions. Where a developer trains a model on a dataset collected by another party, each may be considered an independent controller. Conversely, where the parties jointly determine the purposes and means of collection, they will be deemed joint controllers. The guidance sets out core principles for such processing, including: lawfulness of the data sources and purposes; purpose limitation; transparency (achieved through a public privacy policy describing the data sources and automated tools used, particularly where directly informing data subjects is impossible or would require disproportionate effort); data minimization (for example, by using synthetic data, filtering out certain categories of data, and excluding websites that prohibit scraping or that by their nature contain sensitive information); and ensuring that data is collected from reliable sources, is of adequate quality, and is time-stamped.

The EDPB notes that scraping necessitates one of the six lawful bases under Article 6 of the GDPR. Given that consent is often impracticable in this context, the most common basis invoked is the controller’s legitimate interests. The controller must demonstrate a genuine and clearly articulated legitimate interest, establish that the processing is necessary, and confirm that no less privacy-intrusive alternative is available. Furthermore, the controller must ensure that its interest outweighs the impact on data subjects, taking into account the nature of the data, the context of collection, the potential effects, and data subjects’ reasonable expectations. Where the balancing test does not clearly favor the controller, it must implement risk-mitigation measures such as avoiding sensitive data, limiting collection to publicly available information, limiting the duration of collection, enhancing transparency, and deleting or anonymizing data that is no longer needed. Finally, the guidance clarifies that processing sensitive data requires, in addition to a lawful basis, an applicable exception under the GDPR. However, the EDPB acknowledges that incidental and unintended processing of sensitive data solely within the context of AI training may be permissible, provided the controller implements robust risk-mitigation measures throughout the data’s lifecycle, from collection through the AI system’s deployment and use.

Click here to read the EDPB’s announcement of Guidelines 02/2026 on anonymization and Guidelines 03/2026 on web scraping in the context of generative AI.

MEDIA HIGHLIGHTS