Draft guidance · GDPR + generative AI

The EDPB has put web scraping for generative AI under a more specific GDPR compliance lens.

Guidelines 03/2026 address legal basis, transparency, special-category data, accuracy and data minimisation when web scraping involves personal data processing. The consultation closes on 30 October 2026.

Published and legally reviewed: 5 October 2026 · Reviewed by Constantin Razvan Gospodin.

Direct answer

This is draft EDPB guidance under public consultation — not a new AI Act obligation.

The European Data Protection Board adopted Guidelines 03/2026 on 8 July 2026 and opened them for feedback until 30 October 2026. The guidance explains how the GDPR applies when web scraping used for generative-AI development involves processing personal data. It does not amend the GDPR or the EU AI Act, and the text may change after consultation.

What the EDPB is clarifying

Scraping public web content does not remove GDPR analysis when personal data is involved.

Personal-data processing

The EDPB states that GDPR applies where scraping includes operations such as collecting, storing, organising or retrieving personal data. Public availability is therefore not, by itself, a GDPR exemption.

Legal basis

The draft gives further guidance and examples on legitimate interests in the specific context of web scraping for AI training. Controllers still need a fact-specific Article 6 analysis rather than a generic “AI training” justification.

Transparency

Purpose limitation and transparency require particular attention. Depending on the processing design and the applicable GDPR conditions, individual notification may in some cases be impossible or involve disproportionate effort; that does not eliminate the need to document the legal analysis.

Special-category data

Where Article 9 data is involved, an Article 6 lawful basis is not enough. The EDPB recalls that an Article 9(2) condition is also required and that there is no general exemption for incidental collection.

Data engineering implications

Accuracy and data minimisation now need to be visible in the scraping pipeline.

The EDPB recommends using reliable sources, recording when data was scraped, validating data before AI training and implementing measures to minimise personal data. For governance teams, those are not abstract privacy principles: they point to evidence that should be built into source selection, ingestion, filtering, provenance and training-data controls.

Operational checklist

What GenAI teams should review now

  • Map every scraping source and identify whether personal data is collected or retained.
  • Document the Article 6 legal basis and the actual purpose of the processing.
  • Test whether Article 9 special-category data can enter the pipeline and what controls prevent or limit it.
  • Document the transparency analysis, including any reliance on the GDPR exceptions to individual information duties.
  • Record source reliability, scraping timestamps, validation steps and provenance.
  • Apply collection filters and retention controls that support data minimisation.
  • Keep data-subject-rights, deletion and correction workflows connected to the training-data lifecycle where the GDPR applies.

EU AI Act + GDPR

Do not collapse copyright, AI Act and GDPR reviews into one test.

For a GPAI provider, AI Act Article 53 copyright-policy and training-content-summary duties can sit alongside GDPR obligations where personal data is processed. The legal questions are distinct: a copyright-compliant source is not automatically GDPR-compliant, and GDPR compliance does not resolve the AI Act provider-role or GPAI analysis.

Read the EU AI Act + GDPR framework →

Review GPAI provider obligations →

What remains open

The consultation is still live.

Comments are open until 30 October 2026 at 23:59 CET. Organisations relying on large-scale web scraping for model development should treat the current text as draft supervisory guidance and track the final version rather than freezing controls around language that may still change.