r/gdpr • u/Ok_King6068 • 15h ago
EU 🇪🇺 Meta's AI crawler hit our site 741,900 times last month. Our DPA says we can barely scrape anything. Who are these rules actually for?
I run a large website for a European SME and I looked into where European regulators stand on web scraping. Honestly it surprised me how strict it all is.
The Dutch privacy regulator (AP) published scraping guidance in 2024. Short version: scraping almost always involves personal data, so GDPR applies even if the data is public. Legitimate interest is basically the only legal ground you can use, and the bar is so high that most commercial scraping is simply not allowed. Italy went even further in May 2024, their regulator told website owners to actively defend themselves against AI scrapers with CAPTCHAs and rate limiting. The UK ICO said in December 2024 that scraping for AI is possible in theory, but developers need to be way more transparent and should ask themselves if they can license the data instead. France followed in June 2025 with strict conditions. And the EDPB published draft guidelines on scraping for AI training in July 2026, also strict: robots.txt counts against you in the assessment, and no exception for special category data.
So those are the rules. Now what actually happens. Meta had to pause AI training on EU user posts in June 2024 after pressure from noyb and the Irish DPC. They resumed in May 2025 with an opt out model, noyb says that still violates GDPR, case is ongoing. But that fight was only about Meta's own users. For everyone else's content there was no pause at all. Meta-externalagent, the crawler that Meta itself describes as "downloads website content to include in datasets used for training AI models such as LLMs", visited our website 741,900 times last month. For comparison, Googlebot did 340,100 visits in the same month. And Googlebot at least sends us traffic back. The AI crawler that gives us nothing hits us more than twice as hard. Meanwhile Cloudflare accused Perplexity last year of using stealth crawlers to get around no-crawl rules, and now blocks AI crawlers by default. That says enough about how normal this has become.
To be clear, I actually think the strict rules make sense, they also protect businesses like ours. But right now the result is: European companies read the guidance and don't scrape, while big tech scrapes everything and deals with the lawyers later.
So my question: do you expect regulators to actually go after the big scrapers once the EDPB guidelines are final? Or will it stay like this? Because so far I see a lot of guidance and very little enforcement.