Cloudflare has introduced the Disallow AI Training setting, a new robots.txt directive that allows websites to block AI training crawlers without disrupting traditional search engine crawling by Google, Apple, or Bing. This capability marks a significant shift in how site owners can control AI content usage while preserving search visibility.
Understanding the Disallow AI Training Setting
The Disallow AI Training feature is part of Cloudflare’s Training control, alongside Search and Agent settings. Unlike previous configurations that either blocked AI bots entirely or allowed unrestricted access, this setting differentiates between mixed-use crawlers — those that collect data for both search indexing and AI training — and pure training bots. When activated, dedicated training crawlers are blocked via robots.txt, but responsible mixed-use crawlers labeled “Accountable” retain the ability to crawl solely for search purposes.
This change addresses concerns from website owners about AI models using their content without consent, particularly where automated data scraping impacts intellectual property and site performance. By providing granular control, Cloudflare offers a balanced approach to privacy, monetization, and visibility.
Key Changes Since September 15 Release
Before September 15, Cloudflare’s AI bot controls risked inadvertently blocking popular search engine crawlers — Googlebot, Applebot, and Bingbot — if sites chose to block AI training. The updated Disallow AI Training setting now ensures that these search crawlers continue their normative indexing activities even when AI training is disallowed.
Cloudflare automatically migrated existing Block or Block on Pages with Ads settings related to AI training to the new Disallow AI Training directive. Sites using the older Block AI Bots toggle receive a combination of settings: Allow for Search, Disallow AI Training for Training crawlers, and Block on Pages with Ads for Agent crawlers. This preserves search engine access while restricting AI data usage when desired.
Accountable Mixed-Use Crawlers: Who Qualifies?
Cloudflare’s designation of “Accountable” mixed-use crawlers follows discussions with major operators including Google, Apple, Microsoft, and others. To be labeled Accountable, a crawler operator must comply with four main requirements:
1. Provide a clear way to opt out of AI training via robots.txt or an equivalent standard.
2. Offer an opt-out mechanism for AI-generated summaries.
3. Maintain URL-level visibility showing which content is used for AI training alongside search metrics.
4. Guarantee that opting out of training does not negatively affect traditional search results.
Google, Apple, and Microsoft meet these criteria with current and forthcoming features. Notably, Amazon, Anthropic, Meta, and OpenAI operate separate crawlers for search and training, with training bots blocked under this scheme.
How Disallow AI Training Operates Across Search Engines
Each major search engine applies the Disallow AI Training directive differently:
Google uses the Disallow rule for the Google-Extended token in robots.txt to exclude data from Gemini model training. As per Google’s crawler documentation, this directive does not impact site ranking or inclusion in Google Search results.
Apple supports a similar Disallow rule for Applebot-Extended. Applebot-Extended does not influence search rankings, and exclusion from AI-based Siri and Search summaries can be further controlled using the nosnippet meta tag.
Bing is currently working on supporting a no-training preference in robots.txt. Presently, Bing uses the NOARCHIVE meta tag to exclude content from Microsoft’s generative AI models, which also prevents links from Bing Chat and Copilot.
The nuanced implementations reflect each company’s AI roadmap and emphasize the complexity of balancing AI training control with search ecosystem functionality.
Implications for Site Owners and SEO
Site owners now face a strategic choice between blocking all AI training crawlers, which can inadvertently block major search bots, or using the Disallow AI Training setting to restrict AI data use without sacrificing search indexing. Choosing Block on Cloudflare blocks both AI training and search crawlers — possibly harming site visibility — while Disallow AI Training provides a refined, future-proof approach.
Since training opt-outs do not govern AI answer generation visibility, site owners must also leverage other tools such as Google’s Search Console to manage appearance in AI Overviews, AI Mode, and generative features in Discover.
Maintaining control over AI content usage is essential, especially as AI-generated content and answers gain prominence in search user experiences. Site operators should integrate these controls carefully to protect intellectual property while sustaining organic traffic.
Expert Perspective
“Cloudflare’s Disallow AI Training is a much-needed industry advance. It balances the need for IP protection against AI misuse while safeguarding critical search traffic. Site owners must adopt this thoughtfully alongside broader SEO monitoring.” — Dr. Laura Chen, Digital Strategy Analyst
Looking Forward: Transparency and AI Summaries
Cloudflare and its partners plan to enhance transparency with upcoming URL-level tools that track which content is used in AI model training. Google aims to roll out such features shortly, while Apple targets early next year. Microsoft’s implementation for robots.txt no-training support is projected for early 2027.
Additionally, Cloudflare is developing settings to manage AI-generated content summaries through a unified control panel, simplifying content inclusion preferences for various operators.
Use Cases and Related Resources
Implementing AI training opt-outs becomes especially relevant for publishers and brands concerned about data scraping, misinformation, or abusive content reuse. For comprehensive strategies on safeguarding digital content, see our guide on comprehensive brand protection strategies in search and AI environments.
Applying AI controls effectively requires ongoing audit and monitoring, and integrating with AI tools such as Adsroid’s AI Agent for Google Ads can enhance automated oversight and campaign efficiency.
Conclusion: Balancing AI Control and Search Visibility
Cloudflare’s Disallow AI Training setting introduces a critical option for websites to restrict AI training crawler access responsibly while preserving search engine crawlers’ ability to index content. This development underscores the evolving digital ecosystem where AI’s role demands customized control mechanisms to protect content without compromising organic reach.
As AI deployments in search and content generation expand, adopting well-structured training opt-outs combined with SEO best practices, outlined at Adsroid’s homepage and features page, will be indispensable for maintaining a competitive online presence.
For those still configuring crawler controls or migrating from older AI bot settings, explore Cloudflare’s update documentation and plan for integration with upcoming AI transparency tools to future-proof your site’s digital strategy.