On 8 July 2026, during its latest plenary, the European Data Protection Board (EDPB) adopted guidelines on anonymisation and guidelines on web scraping in the context of generative AI, together with the final version of its guidelines on the processing of personal data through blockchain technologies. The two new documents are presented as separate texts, yet they govern the two ends of the same life cycle: the moment data enters a model, harvested at scale from the web, and the moment it claims to leave the scope of the law because it no longer relates to anyone. A single question holds them together, and it is the hardest question in data protection law: when does information stop being personal data, and who has the power to decide.
Anonymous data is not a quality of the information, it is a relationship
The definition underpinning the anonymisation guidelines is as simple as it is demanding: data are anonymous if they do not relate to an identified or identifiable natural person, and whether this is the case may vary from one entity to another. The qualification is not stamped onto the data like a chemical property; it arises from the relationship between the information and whoever holds it, with the means available to them.
Information can relate to an individual because of its content, its purpose or its effect, and the Board warns that such a link may not be immediately obvious and could require further analysis. An individual is identified or identifiable where they can be distinguished from others in a specific context, using means reasonably likely to be used, in a way that makes it possible to treat them differently; whether those means are reasonably likely to be used depends on the perspective of the relevant entity and must be assessed in light of all objective factors. On this point the guidelines take account of the judgment of the Court of Justice of the European Union in case C-413/23 P, EDPS v SRB, of 4 September 2025, alongside the rest of the Court’s case law.
Anyone who followed the debate on the new definition of personal data proposed by the Digital Omnibus will recognise what is at stake: moving the identifiability threshold even slightly moves the outer boundary of the GDPR and, with it, the reach of the rights of data subjects.
Three criteria and two routes that do not lead to the same place
In operational terms, the EDPB puts forward a framework built on three criteria: no record isolation, no linkage and no inference. Where all three are met, the data can safely be considered anonymous; where even one is not satisfied, further analysis is needed before concluding that anonymisation has succeeded.
The most interesting part, however, concerns how that framework is applied, because the Board allows two approaches. The first, described as contextual, assesses the differences in capabilities between the various entities that might identify the individual, and reflects the full nuances of the legal standard for anonymisation. The second, described as simplified, chooses not to take those differences into account: it is more convenient and offers greater confidence in the result, but it can go beyond the legal standard, leading a controller to treat as non-anonymous data that would in fact be anonymous for some relevant entities.
The admission is remarkable, because it openly accepts that two diligent controllers can reach divergent yet defensible conclusions on the very same dataset. The practical consequence is that the choice of approach becomes a compliance decision in its own right, one to be reasoned and documented, exactly as happens with the assessments that feed into the impact assessment template drawn up by the EDPB itself.
Scraping: no safe harbour at the model’s entrance
If anonymisation governs the exit, web scraping governs the entrance. The Board describes it as a large-scale automated data extraction process that often operates without individuals being aware, and which may pose significant risks to the protection of their personal data. The premise is unambiguous: the GDPR applies to web scraping when it includes personal data processing operations such as collection, storage, organisation and retrieval.
Concrete guidance follows. Particular attention must be paid to the purpose limitation principle and to the transparency principle, on the understanding that, depending on how the processing is precisely designed, the controller might not have to inform individuals personally where doing so proves impossible or would require excessive effort. On accuracy, the EDPB recommends scraping data only from reliable sources, recording the timestamp and validating the data before using them in training, and it advises on further measures to comply with data minimisation. As for legitimate interest, the guidelines do not start from scratch: they build on the Opinion 28/2024 on AI models and bring its criteria down into the specific context of scraping for AI training, with clarifications and examples.
Special categories and the limits of the reasoning in GC and Others
The passage most likely to be debated concerns special categories of personal data. Processing them is in principle prohibited: where scraping involves such data, both a lawful basis under Article 6 and an exception under Article 9(2) of the GDPR are required. The Board accepts that the reasoning of the Court of Justice in GC and Others, case C-136/17, may be relevant to the incidental or residual collection of special categories of data, provided that the controller acts within the “framework of their responsibilities, powers, and capabilities” and implements appropriate technical and organisational measures to prevent the collection and dissemination of such data.
The door, however, is left ajar rather than thrown open: there is no general exemption from the requirements of Article 9, and each case must be assessed individually to determine whether the Court’s reasoning applies. This position strikes at the root of a familiar developer argument, namely that the statistical inevitability of capturing sensitive data within a corpus of billions of pages makes the processing tolerable in itself. Anyone who has grappled with the issue in fields where special categories are the rule rather than the exception, such as the secondary use of health data under the European Health Data Space, knows how narrow that passage is.
The thread running through both texts, and what remains open until 30 October
Read in sequence, the two documents tell the same underlying story. The perimeter of the GDPR is not drawn by an objective property of the data, but by a relational and contextual judgment that the controller must make: at the entrance, when deciding which sources to draw on, which data to exclude and which legal basis to rely on; at the exit, when deciding whether the outcome of the anonymisation process holds up against the three criteria, and under which of the two approaches it should be measured. In both cases the decision falls first to whoever processes the data, and in both cases the price of that autonomy is the burden of demonstrating the soundness of the judgment before a supervisory authority that can reconstruct it after the fact.
At the same plenary, the EDPB adopted the final version of its guidelines on blockchain technologies, accompanied by a report on the outcome of the public consultation and by a track changes version of the text. The anonymisation and web scraping guidelines, by contrast, remain open to public consultation until 30 October 2026: what we read today is not yet the final word, and the scope for reshaping the most contested passages is still there.
In light of the above, one may ask whether a system that leaves controllers to choose between a contextual and a simplified approach, while also asking them to determine case by case when the collection of sensitive data is truly incidental, ends up shifting onto operators the cost of a definition the legislator was unwilling to settle, and whether European businesses are equipped today to bear it.




