Cloudflare has introduced new controls that allow website owners to prevent their content from being used for artificial intelligence training while continuing to remain discoverable through traditional search engines.
The company launched a "Disallow AI Training" setting on September 15 as part of a broader overhaul of how it handles automated crawlers. The move is designed to address mixed-use crawlers that can use the same bot for conventional search indexing as well as collecting content for AI model training.
Previously, blocking such crawlers could create a trade-off for publishers and other website operators. Restricting a bot from accessing content for AI training could also prevent it from indexing the same pages for search, potentially affecting discoverability.
Cloudflare's new framework separates crawler activity into three categories: Search, Training and Agent. Search covers crawling used to index content and help users find it, Training covers content collected for training or fine-tuning AI models, while Agent refers to automated systems accessing websites in real time on behalf of users.
Website owners can now set different preferences for these uses. Cloudflare said Apple, Google and Microsoft either honour or have committed to honouring its new AI training preference within specified timelines. This means eligible mixed-use crawlers can continue accessing content for search while respecting a publisher's request not to use that material for AI training.
The company has also introduced an "Accountable" designation for crawler operators that provide website owners with meaningful controls over how their content is used. Cloudflare said AI summaries are the next area it intends to address, with designated operators required to provide mechanisms allowing publishers to opt out of their content being used in AI-generated summaries.
For new websites carrying advertising, Cloudflare's recommended settings allow Search while disallowing AI Training and blocking Agent activity on pages displaying ads. Websites without advertising can choose different settings, and customers remain able to change the controls through Cloudflare's dashboard.
The distinction is particularly relevant for publishers and digital businesses whose commercial models depend on attracting readers through search, advertising or subscriptions. AI crawlers have created questions around whether publishers can maintain online visibility while limiting how their material is reused to develop AI systems.
Cloudflare has been expanding its tools for managing this traffic as AI systems increasingly interact with web content. Its existing controls include AI Crawl Control and automated robots.txt management, although Cloudflare notes that robots.txt directives express preferences rather than technically preventing access unless additional blocking controls are used.
The latest changes give website operators more granular control over whether automated systems can search, train on or interact with their content, rather than treating all crawler activity as a single category.
Disclaimer: This article may include information derived from interviews, press releases, public statements, research, company communications and other publicly available or third-party sources. Such material may be summarised, paraphrased or contextualised for journalistic and editorial purposes. All rights in third-party content remain with their respective owners.