Cloudflare’s New Disallow AI Training Setting and Its Search Impact

Cloudflare's New Disallow AI Training Setting and Its Search Impact
Cloudflare's Disallow AI Training setting blocks AI training crawlers without impacting Google, Apple, and Bing's search crawling, redefining how sites control AI content use and search visibility.

Summarize with AI

Connect Claude to your Ad Accounts in less than 5mn

Discover the most powerful advertising MCP and unlock 140+ tools to analyze, optimize and manage your campaigns with AI.

Cloudflare has introduced the Disallow AI Training setting, a new robots.txt directive that allows websites to block AI training crawlers without disrupting traditional search engine crawling by Google, Apple, or Bing. This capability marks a significant shift in how site owners can control AI content usage while preserving search visibility.

Understanding the Disallow AI Training Setting

The Disallow AI Training feature is part of Cloudflare’s Training control, alongside Search and Agent settings. Unlike previous configurations that either blocked AI bots entirely or allowed unrestricted access, this setting differentiates between mixed-use crawlers — those that collect data for both search indexing and AI training — and pure training bots. When activated, dedicated training crawlers are blocked via robots.txt, but responsible mixed-use crawlers labeled “Accountable” retain the ability to crawl solely for search purposes.

This change addresses concerns from website owners about AI models using their content without consent, particularly where automated data scraping impacts intellectual property and site performance. By providing granular control, Cloudflare offers a balanced approach to privacy, monetization, and visibility.

Key Changes Since September 15 Release

Before September 15, Cloudflare’s AI bot controls risked inadvertently blocking popular search engine crawlers — Googlebot, Applebot, and Bingbot — if sites chose to block AI training. The updated Disallow AI Training setting now ensures that these search crawlers continue their normative indexing activities even when AI training is disallowed.

Cloudflare automatically migrated existing Block or Block on Pages with Ads settings related to AI training to the new Disallow AI Training directive. Sites using the older Block AI Bots toggle receive a combination of settings: Allow for Search, Disallow AI Training for Training crawlers, and Block on Pages with Ads for Agent crawlers. This preserves search engine access while restricting AI data usage when desired.

Accountable Mixed-Use Crawlers: Who Qualifies?

Cloudflare’s designation of “Accountable” mixed-use crawlers follows discussions with major operators including Google, Apple, Microsoft, and others. To be labeled Accountable, a crawler operator must comply with four main requirements:

1. Provide a clear way to opt out of AI training via robots.txt or an equivalent standard.
2. Offer an opt-out mechanism for AI-generated summaries.
3. Maintain URL-level visibility showing which content is used for AI training alongside search metrics.
4. Guarantee that opting out of training does not negatively affect traditional search results.

Google, Apple, and Microsoft meet these criteria with current and forthcoming features. Notably, Amazon, Anthropic, Meta, and OpenAI operate separate crawlers for search and training, with training bots blocked under this scheme.

How Disallow AI Training Operates Across Search Engines

Each major search engine applies the Disallow AI Training directive differently:

Google uses the Disallow rule for the Google-Extended token in robots.txt to exclude data from Gemini model training. As per Google’s crawler documentation, this directive does not impact site ranking or inclusion in Google Search results.
Apple supports a similar Disallow rule for Applebot-Extended. Applebot-Extended does not influence search rankings, and exclusion from AI-based Siri and Search summaries can be further controlled using the nosnippet meta tag.
Bing is currently working on supporting a no-training preference in robots.txt. Presently, Bing uses the NOARCHIVE meta tag to exclude content from Microsoft’s generative AI models, which also prevents links from Bing Chat and Copilot.

The nuanced implementations reflect each company’s AI roadmap and emphasize the complexity of balancing AI training control with search ecosystem functionality.

Implications for Site Owners and SEO

Site owners now face a strategic choice between blocking all AI training crawlers, which can inadvertently block major search bots, or using the Disallow AI Training setting to restrict AI data use without sacrificing search indexing. Choosing Block on Cloudflare blocks both AI training and search crawlers — possibly harming site visibility — while Disallow AI Training provides a refined, future-proof approach.

Since training opt-outs do not govern AI answer generation visibility, site owners must also leverage other tools such as Google’s Search Console to manage appearance in AI Overviews, AI Mode, and generative features in Discover.

Maintaining control over AI content usage is essential, especially as AI-generated content and answers gain prominence in search user experiences. Site operators should integrate these controls carefully to protect intellectual property while sustaining organic traffic.

Expert Perspective

“Cloudflare’s Disallow AI Training is a much-needed industry advance. It balances the need for IP protection against AI misuse while safeguarding critical search traffic. Site owners must adopt this thoughtfully alongside broader SEO monitoring.” — Dr. Laura Chen, Digital Strategy Analyst

Get Alerts When Competitors Launch New Ads

Ad Radar automatically monitors your competitors across Google, Bing and Meta. Get alerted when a new ad appears for your tracked keywords, so you can spot new offers, messaging and opportunities without constantly checking.

Looking Forward: Transparency and AI Summaries

Cloudflare and its partners plan to enhance transparency with upcoming URL-level tools that track which content is used in AI model training. Google aims to roll out such features shortly, while Apple targets early next year. Microsoft’s implementation for robots.txt no-training support is projected for early 2027.

Additionally, Cloudflare is developing settings to manage AI-generated content summaries through a unified control panel, simplifying content inclusion preferences for various operators.

Use Cases and Related Resources

Implementing AI training opt-outs becomes especially relevant for publishers and brands concerned about data scraping, misinformation, or abusive content reuse. For comprehensive strategies on safeguarding digital content, see our guide on comprehensive brand protection strategies in search and AI environments.

Applying AI controls effectively requires ongoing audit and monitoring, and integrating with AI tools such as Adsroid’s AI Agent for Google Ads can enhance automated oversight and campaign efficiency.

Turn On Copilot. Let AI Optimize Your Ads 24/7.

Your AI agent works in the background, continuously watching your campaigns and finding ways to improve them. When it spots an opportunity, it tells you what to do, and you simply approve the action.

Conclusion: Balancing AI Control and Search Visibility

Cloudflare’s Disallow AI Training setting introduces a critical option for websites to restrict AI training crawler access responsibly while preserving search engine crawlers’ ability to index content. This development underscores the evolving digital ecosystem where AI’s role demands customized control mechanisms to protect content without compromising organic reach.

As AI deployments in search and content generation expand, adopting well-structured training opt-outs combined with SEO best practices, outlined at Adsroid’s homepage and features page, will be indispensable for maintaining a competitive online presence.

For those still configuring crawler controls or migrating from older AI bot settings, explore Cloudflare’s update documentation and plan for integration with upcoming AI transparency tools to future-proof your site’s digital strategy.

Share the post

X
Facebook
LinkedIn

About the author

Picture of Danny Da Rocha - Founder of Adsroid
Danny Da Rocha - Founder of Adsroid
Danny Da Rocha is a digital marketing and automation expert with over 10 years of experience at the intersection of performance advertising, AI, and large-scale automation. He has designed and deployed advanced systems combining Google Ads, data pipelines, and AI-driven decision-making for startups, agencies, and large advertisers. His work has been recognized through multiple industry distinctions for innovation in marketing automation and AI-powered advertising systems. Danny focuses on building practical AI tools that augment human decision-making rather than replacing it.

Table of Contents

Your Google and Meta Ads on Autopilot

Let AI handle the work.

Adsroid analyzes your campaigns, finds opportunities and takes action to improve performance, while you stay in control.

Latest posts

AI Ad Automation Statistics 2026: Key Data on Autonomous Advertising

Key AI ad automation statistics and autonomous advertising data for 2026. Explore market size, adoption rates, platform trends, and what the numbers mean for paid media teams.

How to Ask an AI Agent ‘Why Did My Traffic Drop?’ Using Connected Ads and Analytics Data

Yes, an AI agent can diagnose why your website traffic dropped. This article walks through a real cross-channel diagnostic conversation combining Google Ads, GA4, and Search Console data.

How to Track Competitor Ads Automatically: A Step-by-Step Setup Guide

Learn how to automatically track competitor ads across Google, Bing, and Meta with a practical step-by-step setup guide covering tools, workflows, alerts, and common mistakes to avoid.