September 15, 2026: Cloudflare's crawler split and how to configure Search, Agent and Training

Cloudflare's September 15, 2026 defaults block Training and Agent bots on ad pages, and can block Googlebot too. Configure the three classes deliberately.

Read time
14 min
Word count
2K
Sections
11
FAQs
8
Share
Editorial graphic on Cloudflare's Search, Agent and Training AI crawler classes with a Sept 15, 2026 deadline
Cloudflare splits AI crawlers into Search, Agent and Training, with new defaults from September 15, 2026.
On this page · 11 sections
  1. What changed on July 1, 2026
  2. The September 15 default flip, in plain terms
  3. The Googlebot trap
  4. Why Cloudflare is forcing the question: the crawl economics
  5. How to configure the three classes, step by step
  6. Which setting each kind of site should pick
  7. India-specific considerations
  8. Your checklist before September 15, 2026
  9. FAQ
  10. How eCorpIT can help
  11. References

Summary. On July 1, 2026 Cloudflare gave every customer, including the $0 Free tier, three separate switches for AI traffic: Search, Agent, and Training. On September 15, 2026 the defaults flip. For new domains, Training and Agent crawlers get blocked on any page that shows ads, while Search stays allowed. The catch sits in one line of Cloudflare's own announcement: multi-purpose crawlers such as Googlebot, Applebot, and BingBot are judged by all of their behaviors, so a site that blocks Training also blocks those search bots unless it opts out. Bots already passed humans on the network at 57.5% to 42.5% by mid-2026, and Cloudflare Radar data put Anthropic's ClaudeBot at 11,122 pages crawled per referral against Googlebot's 5 to 1. This guide shows publishers, SEO and GEO leads exactly which of the three switches to set, how to keep Google Search access, and what to change before the deadline.

Cloudflare sits in front of more than 20% of web domains, so a default change on its network is a change most site owners inherit whether they read the release notes or not. The company framed the update as its second "Content Independence Day," one year after it first let site owners block AI training crawlers. Matthew Prince, co-founder and CEO of Cloudflare, tied the move to raw traffic numbers: "Now that the majority of traffic on the Internet is non-human, we must go further and act faster so that a sustainable ecosystem can emerge." The practical result is that a setting you never touched now has three positions instead of one, and the wrong position can pull you out of Google.

What changed on July 1, 2026

Cloudflare replaced the single "Block AI bots" toggle with a taxonomy built on what a bot does rather than whether it counts as "AI." Its reasoning is that "AI can be in anything" now that Google Search answers questions directly on the results page, so a label like "AI bot" ages badly. The three classes it wants every customer to manage are Search, Agent, and Training. All three are configurable today on every plan, from the Free tier at $0 to Business at $250 per month billed monthly, and Enterprise.

The definitions matter, because your setting keys off them:

Class What the bot does Default on ad pages from Sept 15, 2026
Search Indexes your content to answer questions later; expected to send referral traffic or pay Allowed
Agent Acts in real time for a person, for example ChatGPT-User or Gemini and Claude driving a browser Blocked
Training Pulls content to train or fine-tune a model; content is absorbed permanently Blocked
Multi-purpose (Search + Training) One crawler that both indexes and trains, for example Googlebot Blocked if you block Training
SEO, Ads Verification, Feed Fetching Site audits, ad checks, RSS and podcast readers Managed separately, not part of the three-way default

Cloudflare also launched supporting pieces the same day: BotBase, a searchable directory of tracked bots for Enterprise Bot Management; a Monetization Gateway that lets sites charge bots for access and settle in stablecoins over the x402 protocol; and a content-use signal covering what a bot may keep after it crawls. Those are context. The setting that will change traffic for most sites is the three-way default.

The September 15 default flip, in plain terms

Two things happen on September 15, 2026, and only for sites that have not set their own preference.

First, for all new domains onboarding to Cloudflare, Training and Agent are blocked by default on pages that display ads, and Search stays allowed. Cloudflare's logic is that an ad signals a page meant for a human to land on and see, so it keeps away bots that consume the page without sending a person, and it keeps Search because Search is the class that funnels visitors back.

Second, and this is the part that trips people up, multi-purpose crawlers that combine Search with Training are allowed or blocked by all of their behaviors, enforced by the most restrictive rule that applies. Cloudflare names the affected bots directly: Googlebot, Applebot, and BingBot will be blocked for any customer who has chosen to block Training, whether through the new controls or the legacy "Block AI bots" service. If you are an existing customer who once clicked "Block AI bots," that legacy choice now reaches further than it did when you set it.

You can opt out. Cloudflare lets any site owner mark, in Security settings at any point before September 15, that they want no changes to Training crawlers that also crawl for Search. That single opt-out is the difference between keeping Googlebot and losing it, and it is the first thing most ad-supported publishers should do.

The Googlebot trap

Google does not run a separate training-only crawler that you can block in isolation. Its flagship Googlebot crawls for Search, and that same index feeds AI Overviews and AI Mode. Google offers Google-Extended as an opt-out for training its Gemini apps and Vertex AI, and using it does not remove you from Google Search. But Google-Extended does not solve the Cloudflare default, because Cloudflare classifies Googlebot itself as multi-purpose. Block Training at the Cloudflare layer without opting out, and you can block the crawler that keeps you in Google.

This is why the same setting helps one site and hurts another. A newsroom that wants payment for training data and an ecommerce store that just wants to stay in Google need opposite choices. The table below shows how the major crawlers land under the new classification, based on Cloudflare's announcement and reporting from TechCrunch and Search Engine Journal.

Crawler Operator How Cloudflare treats it Blocked if you block Training on ad pages?
Googlebot Google Multi-purpose: Search plus training signals Yes, unless you opt out
Applebot Apple Multi-purpose Yes, unless you opt out
BingBot Microsoft Multi-purpose Yes, unless you opt out
GPTBot OpenAI Training Yes
ChatGPT-User OpenAI Agent Yes (Agent default)
ClaudeBot Anthropic Training Yes
Google-Extended Google Training opt-out token, not a blockable search bot Does not restore Search access on its own

The design goal is transparency: Cloudflare wants operators to run separate crawlers for separate jobs, and it argues that the ones that already do, such as OpenAI and Anthropic, are easier for site owners to manage than a single blended bot. Whether you agree with the pressure campaign or not, the operational fact stands. If your revenue depends on Google referrals, you must opt out of the Training default or you risk your own indexing.

Why Cloudflare is forcing the question: the crawl economics

The reason a CDN is rewriting crawler defaults at all is that the old bargain broke. For thirty years the deal was simple: crawlers took your pages and sent back readers. AI crawling inverted that. Cloudflare Radar data reported through mid-2026 quantifies the gap, and the numbers are the strongest argument in the whole debate.

Operator or bot Crawl-to-referral ratio Period (Cloudflare Radar)
Googlebot (traditional search) about 5 to 1 Mid-2026
OpenAI about 857 to 1 Late May to early June 2026
Anthropic ClaudeBot about 11,122 to 1 Late May to early June 2026
Anthropic ClaudeBot about 23,951 to 1 January to March 2026
All bots vs humans on network 57.5% to 42.5% Mid-2026

Read the ClaudeBot line carefully. Anthropic crawls heavily to train Claude but runs no consumer search product that returns clicks, which is why its ratio dwarfs Google's. The ratio improved from roughly 23,951 to 1 early in 2026 to about 11,122 to 1 by early summer, so the trend is moving, but a five-figure crawl-to-referral gap is still a one-way transfer. Training returns almost nothing, and most AI crawling is training rather than search. That is the case Cloudflare is making with its defaults: keep the class that sends readers, gate the classes that do not.

For context on how little of the click now reaches publishers even when a page does rank, our analysis of ranking versus AI Overview citation data and the keyword selection framework for click survival both show why raw impressions stopped translating into visits in 2026.

How to configure the three classes, step by step

You control all of this from the Cloudflare dashboard, and you can back it with a robots.txt signal. Here is the deliberate setup.

  1. Open your zone in the Cloudflare dashboard and go to Security, then Settings. The AI traffic controls for Search, Agent, and Training are there for every plan, including Free.
  1. Decide per class. Allow Search unless you specifically want to charge for indexing. Choose Agent based on whether you want real-time assistants fetching your pages. Choose Training based on whether you are willing to have your content absorbed into models without payment.
  1. If you run ads and depend on Google, set your preference before September 15 and use the opt-out that keeps Training-plus-Search crawlers, so Googlebot, Applebot, and BingBot stay allowed. Cloudflare exposes this directly in Security settings.
  1. Back the choice in robots.txt with Content Signals. Cloudflare's managed robots.txt writes preferences that AI operators are meant to honor. The new content-use field sets how far a bot may go: immediate (use nothing), reference (index, excerpt, link back, the default), or full (summarize and reproduce).

Cloudflare's managed robots.txt now looks like this:


            User-agent: *
Content-Signal: search=yes,ai-train=no,use=reference
Allow: /
          

A signal is a stated preference, not a hard block; the dashboard setting is what actually allows or denies traffic at the edge. Use both so your intent is legible to operators that respect robots.txt and enforced against those that do not. If you already run bot rules, sanity-check them against your broader bot-detection setup; our note on Cloudflare's session-based bot detection covers how these layers interact.

Which setting each kind of site should pick

There is no universal answer, because the three classes trade discoverability against control. Map your choice to what the site is for.

Site type Suggested setting Why
Ad-monetized publisher or news site Allow Search, block Training, opt out to keep Googlebot Protects training value while staying in Google and AI Overviews
Ecommerce or SaaS with no ads Allow Search and Agent, decide Training case by case Agents completing tasks can help conversion; ads default does not apply
Small blog needing discovery Allow Search, allow or block Training Discoverability matters more than training payment at low traffic
Documentation or support site Allow Search and Agent, allow Training at reference Being cited accurately in AI answers is the goal
Premium or paywalled content Block Training and Agent, consider Monetization Gateway Charge for access rather than give it away

The opt-out to keep Googlebot is the single most important control for anyone whose traffic comes from Google. Set it deliberately rather than inheriting a default that was designed for a different kind of site. For the wider strategy of staying visible as Google shifts users into AI Mode by default, and for how structured, citable content earns AI Overview citations, crawler access is the precondition: an engine cannot cite a page it was blocked from reading.

India-specific considerations

Indian publishers and D2C brands sit in the same trap with an added compliance layer. Ad-monetized Indian news and content sites that rely on Google referral traffic must set the Googlebot opt-out before September 15, 2026, or risk losing organic reach at the same time the market is shifting to AI answers. For Indian SaaS and fintech firms, the Agent class is the one to watch, because real-time assistants acting on a user's behalf increasingly complete tasks inside web apps, and blocking them can break legitimate automated flows.

On the data side, the content-use levels connect to India's Digital Personal Data Protection Act. Any page that exposes personal data should lean toward immediate or reference rather than full, so a model does not reproduce personal information verbatim. Teams building consent and data-handling flows can pair this with our DPDP engineering playbook for Indian startups. The crawler setting and the privacy obligation point the same way: control what leaves your site and how it can be reused.

Your checklist before September 15, 2026

Do these in order. First, confirm whether your key pages display ads, because the default only bites on ad pages. Second, if Google traffic matters, set the opt-out in Security settings that keeps Training-plus-Search crawlers. Third, set Search, Agent, and Training deliberately per the site-type table rather than leaving them on the incoming default. Fourth, enable managed robots.txt with a use=reference signal so your intent is stated as well as enforced. Fifth, if you run premium content, evaluate the Monetization Gateway instead of a flat block. Finally, watch your server logs and Cloudflare analytics for the two weeks after the deadline to catch any crawler you blocked by accident.

FAQ

How eCorpIT can help

eCorpIT is a senior-led engineering and digital marketing organisation in Gurugram that helps publishers and product teams configure crawler policy without losing search or AI-search visibility. We audit your Cloudflare bot settings, set the Search, Agent, and Training classes against your business model, and align robots.txt Content Signals with your DPDP obligations. If the September 15, 2026 deadline affects your traffic, contact our team or read more about our GEO and AEO optimisation service.

References

  1. Cloudflare Blog — Your site, your rules: new AI traffic options for all customers (July 1, 2026)
  1. Cloudflare Press — Cloudflare allows the agentic Internet to flourish: your content, your rules
  1. Cloudflare Blog — Content Independence Day, one year on: building the business model for the agentic Internet
  1. Cloudflare Blog — Making AI search smarter
  1. Cloudflare Blog — Announcing the Monetization Gateway
  1. Cloudflare Blog — Google's AI advantage: why crawler separation is the only path to a fair Internet
  1. Cloudflare Developer Docs — Block AI bots
  1. Cloudflare Developer Docs — BotBase
  1. TechCrunch — Cloudflare's new policy pushes AI companies to pay for publishers' content
  1. Search Engine Journal — Cloudflare's AI crawler rules can block Googlebot
  1. Content Signals — Content Signals specification
  1. SEOmator — GEO data report 2026: crawl-to-refer ratios for AI crawlers and LLM bots
  1. Spendbase — Cloudflare pricing explained: Free, Pro, Business and Enterprise plans

_Last updated: July 23, 2026._

Frequently asked

Quick answers.

01 What exactly changes on September 15, 2026?
For new domains on Cloudflare that have not set a preference, Training and Agent crawlers are blocked by default on pages showing ads, while Search stays allowed. Multi-purpose crawlers that also train, such as Googlebot, get blocked for anyone who blocks Training unless they opt out first.
02 Will this block Google from my site?
It can. Googlebot crawls for Search and training signals in one bot, so Cloudflare treats it as multi-purpose. If you block Training without using Cloudflare's opt-out for Training-plus-Search crawlers, Googlebot is blocked on ad pages. Setting the opt-out in Security settings before the deadline keeps Google access.
03 Do the new controls cost anything?
No. Cloudflare made the Search, Agent, and Training controls available to every plan on July 1, 2026, including the Free tier at $0 per month. Paid tiers such as Pro at $25 and Business at $250 per month add other features, but the three-way AI traffic setting itself is not gated behind a paid plan.
04 What is the difference between Agent and Training bots?
An Agent acts in real time for a person, such as ChatGPT-User fetching a page or Claude driving a browser to finish a task. Training crawlers pull your content to train or fine-tune a model, absorbing it permanently. Cloudflare blocks both by default on ad pages, but for different reasons.
05 Does robots.txt alone stop AI crawlers?
No. A robots.txt Content Signal such as search=yes,ai-train=no,use=reference states your preference, and cooperative operators are meant to honor it. It does not force compliance. The dashboard setting is what allows or denies traffic at Cloudflare's edge, so use the signal and the setting together for both intent and enforcement.
06 How bad is the crawl-to-referral imbalance?
By Cloudflare Radar data reported in mid-2026, traditional Googlebot ran about 5 pages crawled per referral, OpenAI about 857 to 1, and Anthropic's ClaudeBot about 11,122 to 1. ClaudeBot had improved from roughly 23,951 to 1 earlier in 2026. Bots also passed humans on the network at 57.5% to 42.5%.
07 Should Indian publishers do anything different?
Ad-monetized Indian sites that depend on Google referrals should set the Googlebot opt-out before September 15, 2026, to avoid losing reach. On the data side, use immediate or reference content-use levels for pages holding personal data, aligning crawler settings with Digital Personal Data Protection Act obligations on reuse.
08 What if I already clicked "Block AI bots" last year?
That legacy choice now counts as blocking Training under the new rules, so from September 15 it can also block multi-purpose crawlers like Googlebot, Applebot, and BingBot on ad pages. Review your Security settings and opt out for Training-plus-Search crawlers if you want to keep search engine access.

About the author

Manu Shukla

Founder & Director

Founder of eCorpIT. Hands-on engineer leading senior-only delivery for AI apps, custom software, and cloud systems for global clients.

Subscribe

One engineering note a week. No fluff, no spam.

Senior-architect playbooks on AI agents, mobile apps, cloud, security, data, and marketing — delivered every Wednesday.

Past the reading

Read enough. Let's build something.

A senior architect responds in 24 working hours with scope, indicative cost, and a timeline. NDA before any technical conversation.