On this page · 11 sections
- What changed on July 1, 2026
- The September 15 default flip, in plain terms
- The Googlebot trap
- Why Cloudflare is forcing the question: the crawl economics
- How to configure the three classes, step by step
- Which setting each kind of site should pick
- India-specific considerations
- Your checklist before September 15, 2026
- FAQ
- How eCorpIT can help
- References
Summary. On July 1, 2026 Cloudflare gave every customer, including the $0 Free tier, three separate switches for AI traffic: Search, Agent, and Training. On September 15, 2026 the defaults flip. For new domains, Training and Agent crawlers get blocked on any page that shows ads, while Search stays allowed. The catch sits in one line of Cloudflare's own announcement: multi-purpose crawlers such as Googlebot, Applebot, and BingBot are judged by all of their behaviors, so a site that blocks Training also blocks those search bots unless it opts out. Bots already passed humans on the network at 57.5% to 42.5% by mid-2026, and Cloudflare Radar data put Anthropic's ClaudeBot at 11,122 pages crawled per referral against Googlebot's 5 to 1. This guide shows publishers, SEO and GEO leads exactly which of the three switches to set, how to keep Google Search access, and what to change before the deadline.
Cloudflare sits in front of more than 20% of web domains, so a default change on its network is a change most site owners inherit whether they read the release notes or not. The company framed the update as its second "Content Independence Day," one year after it first let site owners block AI training crawlers. Matthew Prince, co-founder and CEO of Cloudflare, tied the move to raw traffic numbers: "Now that the majority of traffic on the Internet is non-human, we must go further and act faster so that a sustainable ecosystem can emerge." The practical result is that a setting you never touched now has three positions instead of one, and the wrong position can pull you out of Google.
What changed on July 1, 2026
Cloudflare replaced the single "Block AI bots" toggle with a taxonomy built on what a bot does rather than whether it counts as "AI." Its reasoning is that "AI can be in anything" now that Google Search answers questions directly on the results page, so a label like "AI bot" ages badly. The three classes it wants every customer to manage are Search, Agent, and Training. All three are configurable today on every plan, from the Free tier at $0 to Business at $250 per month billed monthly, and Enterprise.
The definitions matter, because your setting keys off them:
| Class | What the bot does | Default on ad pages from Sept 15, 2026 |
|---|---|---|
| Search | Indexes your content to answer questions later; expected to send referral traffic or pay | Allowed |
| Agent | Acts in real time for a person, for example ChatGPT-User or Gemini and Claude driving a browser | Blocked |
| Training | Pulls content to train or fine-tune a model; content is absorbed permanently | Blocked |
| Multi-purpose (Search + Training) | One crawler that both indexes and trains, for example Googlebot | Blocked if you block Training |
| SEO, Ads Verification, Feed Fetching | Site audits, ad checks, RSS and podcast readers | Managed separately, not part of the three-way default |
Cloudflare also launched supporting pieces the same day: BotBase, a searchable directory of tracked bots for Enterprise Bot Management; a Monetization Gateway that lets sites charge bots for access and settle in stablecoins over the x402 protocol; and a content-use signal covering what a bot may keep after it crawls. Those are context. The setting that will change traffic for most sites is the three-way default.
The September 15 default flip, in plain terms
Two things happen on September 15, 2026, and only for sites that have not set their own preference.
First, for all new domains onboarding to Cloudflare, Training and Agent are blocked by default on pages that display ads, and Search stays allowed. Cloudflare's logic is that an ad signals a page meant for a human to land on and see, so it keeps away bots that consume the page without sending a person, and it keeps Search because Search is the class that funnels visitors back.
Second, and this is the part that trips people up, multi-purpose crawlers that combine Search with Training are allowed or blocked by all of their behaviors, enforced by the most restrictive rule that applies. Cloudflare names the affected bots directly: Googlebot, Applebot, and BingBot will be blocked for any customer who has chosen to block Training, whether through the new controls or the legacy "Block AI bots" service. If you are an existing customer who once clicked "Block AI bots," that legacy choice now reaches further than it did when you set it.
You can opt out. Cloudflare lets any site owner mark, in Security settings at any point before September 15, that they want no changes to Training crawlers that also crawl for Search. That single opt-out is the difference between keeping Googlebot and losing it, and it is the first thing most ad-supported publishers should do.
The Googlebot trap
Google does not run a separate training-only crawler that you can block in isolation. Its flagship Googlebot crawls for Search, and that same index feeds AI Overviews and AI Mode. Google offers Google-Extended as an opt-out for training its Gemini apps and Vertex AI, and using it does not remove you from Google Search. But Google-Extended does not solve the Cloudflare default, because Cloudflare classifies Googlebot itself as multi-purpose. Block Training at the Cloudflare layer without opting out, and you can block the crawler that keeps you in Google.
This is why the same setting helps one site and hurts another. A newsroom that wants payment for training data and an ecommerce store that just wants to stay in Google need opposite choices. The table below shows how the major crawlers land under the new classification, based on Cloudflare's announcement and reporting from TechCrunch and Search Engine Journal.
| Crawler | Operator | How Cloudflare treats it | Blocked if you block Training on ad pages? |
|---|---|---|---|
| Googlebot | Multi-purpose: Search plus training signals | Yes, unless you opt out | |
| Applebot | Apple | Multi-purpose | Yes, unless you opt out |
| BingBot | Microsoft | Multi-purpose | Yes, unless you opt out |
| GPTBot | OpenAI | Training | Yes |
| ChatGPT-User | OpenAI | Agent | Yes (Agent default) |
| ClaudeBot | Anthropic | Training | Yes |
| Google-Extended | Training opt-out token, not a blockable search bot | Does not restore Search access on its own |
The design goal is transparency: Cloudflare wants operators to run separate crawlers for separate jobs, and it argues that the ones that already do, such as OpenAI and Anthropic, are easier for site owners to manage than a single blended bot. Whether you agree with the pressure campaign or not, the operational fact stands. If your revenue depends on Google referrals, you must opt out of the Training default or you risk your own indexing.
Why Cloudflare is forcing the question: the crawl economics
The reason a CDN is rewriting crawler defaults at all is that the old bargain broke. For thirty years the deal was simple: crawlers took your pages and sent back readers. AI crawling inverted that. Cloudflare Radar data reported through mid-2026 quantifies the gap, and the numbers are the strongest argument in the whole debate.
| Operator or bot | Crawl-to-referral ratio | Period (Cloudflare Radar) |
|---|---|---|
| Googlebot (traditional search) | about 5 to 1 | Mid-2026 |
| OpenAI | about 857 to 1 | Late May to early June 2026 |
| Anthropic ClaudeBot | about 11,122 to 1 | Late May to early June 2026 |
| Anthropic ClaudeBot | about 23,951 to 1 | January to March 2026 |
| All bots vs humans on network | 57.5% to 42.5% | Mid-2026 |
Read the ClaudeBot line carefully. Anthropic crawls heavily to train Claude but runs no consumer search product that returns clicks, which is why its ratio dwarfs Google's. The ratio improved from roughly 23,951 to 1 early in 2026 to about 11,122 to 1 by early summer, so the trend is moving, but a five-figure crawl-to-referral gap is still a one-way transfer. Training returns almost nothing, and most AI crawling is training rather than search. That is the case Cloudflare is making with its defaults: keep the class that sends readers, gate the classes that do not.
For context on how little of the click now reaches publishers even when a page does rank, our analysis of ranking versus AI Overview citation data and the keyword selection framework for click survival both show why raw impressions stopped translating into visits in 2026.
How to configure the three classes, step by step
You control all of this from the Cloudflare dashboard, and you can back it with a robots.txt signal. Here is the deliberate setup.
- Open your zone in the Cloudflare dashboard and go to Security, then Settings. The AI traffic controls for Search, Agent, and Training are there for every plan, including Free.
- Decide per class. Allow Search unless you specifically want to charge for indexing. Choose Agent based on whether you want real-time assistants fetching your pages. Choose Training based on whether you are willing to have your content absorbed into models without payment.
- If you run ads and depend on Google, set your preference before September 15 and use the opt-out that keeps Training-plus-Search crawlers, so Googlebot, Applebot, and BingBot stay allowed. Cloudflare exposes this directly in Security settings.
- Back the choice in robots.txt with Content Signals. Cloudflare's managed robots.txt writes preferences that AI operators are meant to honor. The new content-use field sets how far a bot may go:
immediate(use nothing),reference(index, excerpt, link back, the default), orfull(summarize and reproduce).
Cloudflare's managed robots.txt now looks like this:
User-agent: *
Content-Signal: search=yes,ai-train=no,use=reference
Allow: /
A signal is a stated preference, not a hard block; the dashboard setting is what actually allows or denies traffic at the edge. Use both so your intent is legible to operators that respect robots.txt and enforced against those that do not. If you already run bot rules, sanity-check them against your broader bot-detection setup; our note on Cloudflare's session-based bot detection covers how these layers interact.
Which setting each kind of site should pick
There is no universal answer, because the three classes trade discoverability against control. Map your choice to what the site is for.
| Site type | Suggested setting | Why |
|---|---|---|
| Ad-monetized publisher or news site | Allow Search, block Training, opt out to keep Googlebot | Protects training value while staying in Google and AI Overviews |
| Ecommerce or SaaS with no ads | Allow Search and Agent, decide Training case by case | Agents completing tasks can help conversion; ads default does not apply |
| Small blog needing discovery | Allow Search, allow or block Training | Discoverability matters more than training payment at low traffic |
| Documentation or support site | Allow Search and Agent, allow Training at reference |
Being cited accurately in AI answers is the goal |
| Premium or paywalled content | Block Training and Agent, consider Monetization Gateway | Charge for access rather than give it away |
The opt-out to keep Googlebot is the single most important control for anyone whose traffic comes from Google. Set it deliberately rather than inheriting a default that was designed for a different kind of site. For the wider strategy of staying visible as Google shifts users into AI Mode by default, and for how structured, citable content earns AI Overview citations, crawler access is the precondition: an engine cannot cite a page it was blocked from reading.
India-specific considerations
Indian publishers and D2C brands sit in the same trap with an added compliance layer. Ad-monetized Indian news and content sites that rely on Google referral traffic must set the Googlebot opt-out before September 15, 2026, or risk losing organic reach at the same time the market is shifting to AI answers. For Indian SaaS and fintech firms, the Agent class is the one to watch, because real-time assistants acting on a user's behalf increasingly complete tasks inside web apps, and blocking them can break legitimate automated flows.
On the data side, the content-use levels connect to India's Digital Personal Data Protection Act. Any page that exposes personal data should lean toward immediate or reference rather than full, so a model does not reproduce personal information verbatim. Teams building consent and data-handling flows can pair this with our DPDP engineering playbook for Indian startups. The crawler setting and the privacy obligation point the same way: control what leaves your site and how it can be reused.
Your checklist before September 15, 2026
Do these in order. First, confirm whether your key pages display ads, because the default only bites on ad pages. Second, if Google traffic matters, set the opt-out in Security settings that keeps Training-plus-Search crawlers. Third, set Search, Agent, and Training deliberately per the site-type table rather than leaving them on the incoming default. Fourth, enable managed robots.txt with a use=reference signal so your intent is stated as well as enforced. Fifth, if you run premium content, evaluate the Monetization Gateway instead of a flat block. Finally, watch your server logs and Cloudflare analytics for the two weeks after the deadline to catch any crawler you blocked by accident.
FAQ
How eCorpIT can help
eCorpIT is a senior-led engineering and digital marketing organisation in Gurugram that helps publishers and product teams configure crawler policy without losing search or AI-search visibility. We audit your Cloudflare bot settings, set the Search, Agent, and Training classes against your business model, and align robots.txt Content Signals with your DPDP obligations. If the September 15, 2026 deadline affects your traffic, contact our team or read more about our GEO and AEO optimisation service.
References
- Cloudflare Blog — Your site, your rules: new AI traffic options for all customers (July 1, 2026)
- Cloudflare Press — Cloudflare allows the agentic Internet to flourish: your content, your rules
- Cloudflare Blog — Content Independence Day, one year on: building the business model for the agentic Internet
- Cloudflare Blog — Making AI search smarter
- Cloudflare Blog — Announcing the Monetization Gateway
- Cloudflare Developer Docs — Block AI bots
- Cloudflare Developer Docs — BotBase
- Search Engine Journal — Cloudflare's AI crawler rules can block Googlebot
- Content Signals — Content Signals specification
_Last updated: July 23, 2026._