Have it both ways: stay discoverable in search while disallowing AI training
Curated from Cloudflare Blog
If you manage high-traffic web properties, the tension between search visibility and data sovereignty is no longer theoretical. Cloudflare’s latest update signals a critical shift in how we handle AI crawlers, moving beyond simple blocking toward granular, industry-standard controls. This matters because your site’s architecture now directly impacts your exposure to unauthorized model training. By aligning with major tech players, Cloudflare is formalizing a distinction that many SREs have been hacking together with robots.txt variations. For your infrastructure team, this means auditing your current crawler policies is no longer optional. You need to decide explicitly which agents are permitted to index your content for search versus those that might ingest it for training. The concrete takeaway is to immediately review your WAF and bot management rules to ensure you are not inadvertently allowing broad AI training crawlers while trying to maintain SEO presence.
Cloudflare is giving site owners a way to stay discoverable while disallowing AI training. New controls and an Accountable designation establish a shared model with Apple, Google, and Microsoft.
— Cloudflare Blog