On this page · 12 sections
- What the ratio actually measures
- The caveat Cloudflare published and almost nobody repeats
- The verified twelve-month series
- Referral share is a separate number, and it is small
- What Cloudflare's CEO is actually arguing for
- A crawler policy that holds up
- Measure your own ratio before you copy anyone's policy
- India-specific considerations
- What to do this month
- FAQ
- How eCorpIT can help
- References
Summary. In July 2025, Cloudflare published a ratio that went round the world: Anthropic's crawlers made roughly 70,900 HTML page requests for every referral they sent back, against Mistral at 0.1:1. Twelve months later, Cloudflare Radar's own crawl_refer_ratio endpoint puts Anthropic at 1,917:1 and OpenAI at 251:1 for the full month of July 2026, against 38,744:1 and 1,104:1 in July 2025. That is a 95% fall for Anthropic and a 77% fall for OpenAI in one year. Google sits at 4.7:1 and DuckDuckGo at 2.4:1. Meanwhile chatgpt.com supplied 0.913% of crawler referrals in July 2026 against Google's 88.08%. Most of the ratios still being quoted in blog posts and board decks this month are the 2025 numbers, and the decision they are being used to justify — block the AI crawlers — is being made on stale data.
There is a second problem, and it comes from Cloudflare rather than from anyone misreading Cloudflare. Every published ratio is an overstatement by construction. Here is the verified series, the methodology caveat that nobody quotes, and a policy that survives both.
What the ratio actually measures
Cloudflare defines it precisely in the original July 2025 post by David Belson and Sam Rhea:
"As HTML pages are arguably the most valuable content for these crawlers, the ratios displayed are calculated by dividing the total number of requests from relevant user agents associated with a given search or AI platform where the response was of Content-type: text/html by the total number of requests for HTML content where the Referer header contained a hostname associated with a given search or AI platform."
The ratio is then normalised to a single referral request. So 1,917:1 means 1,917 HTML page requests from Anthropic's crawlers for every one HTML page view that arrived carrying an Anthropic hostname in the Referer header.
Read that definition twice, because the second half is where the number breaks.
The caveat Cloudflare published and almost nobody repeats
From the same post: "traffic referred by Claude's native app does not include a Referer: header, and we believe that the same holds true for traffic generated from other native apps as well. As such, because the referral counts only include traffic from the Web-based tools from these providers, these calculations may overstate the respective ratios, but it is unclear by how much."
That is Cloudflare saying, in its own publication, that the denominator is incomplete.
The consequence is not small. Anthropic, OpenAI and Perplexity all push heavily on desktop and mobile apps. Every citation a user taps inside the Claude desktop app, the ChatGPT iOS app or the Perplexity Android app is a real referral that never appears in the denominator. The crawl side of the fraction is measured accurately; the referral side is measured only for the web client. The ratio is therefore a ceiling, not a measurement, and the more app-heavy a platform gets, the worse the overstatement.
Three further caveats sit in the same methodology note: only text/html responses are counted, multiple user agents per platform are aggregated under one platform name, and referral traffic from Google's own ASN (AS15169) is excluded because prefetch is not human consumption.
None of this makes the ratios useless. It makes them a directional signal about business model, which is exactly how the numbers below should be read.
The verified twelve-month series
These figures come from Cloudflare Radar's bots/crawlers/summary/crawl_refer_ratio endpoint with RATIO normalisation, pulled on 1 August 2026 and published in TechnologyChecker's ChatGPT statistics report by Emma Davies, a data analyst there. All three columns are full calendar months, so they are directly comparable with each other.
| Platform | July 2025 | June 2026 | July 2026 |
|---|---|---|---|
| Anthropic | 38,744:1 | 3,386:1 | 1,917:1 |
| OpenAI | 1,104:1 | 647:1 | 251:1 |
| Perplexity | 195:1 | 207:1 | 289:1 |
| Microsoft | 40.4:1 | 35.3:1 | 36.2:1 |
| ByteDance | 0.9:1 | 10.1:1 | 9.6:1 |
| 5.5:1 | 5.0:1 | 4.7:1 | |
| DuckDuckGo | 0.3:1 | 2.1:1 | 2.4:1 |
Four readings that matter more than the headline.
Anthropic's collapse is the biggest story and the least reported. 38,744:1 to 1,917:1 is a 95% improvement in twelve months. Anthropic did not become a search engine; what changed is that Claude now surfaces and links sources far more often, and its crawl volume settled after the early training-scrape surge. If your crawler policy was written against a five-figure ratio, it was written against a number that no longer exists.
OpenAI improved 77% and is now inside two orders of magnitude of Microsoft. 1,104:1 to 251:1 puts GPTBot's give-back within striking distance of Bing's 36.2:1. The direction of travel is unambiguous.
Perplexity went the wrong way. 195:1 in July 2025, 207:1 in June 2026, 289:1 in July 2026. It is the only major platform in the table whose ratio deteriorated year on year, and it deteriorated 40% in a single month. Perplexity is the platform most often described as citation-friendly. On this metric it is currently getting worse, not better.
DuckDuckGo went from 0.3:1 to 2.4:1 — it now crawls more per referral than it used to, moving toward the others rather than away from them. Google drifted the other way, from 5.5:1 to 4.7:1.
Why the AI ratios improved so fast is worth being clear about, because the reason determines whether the trend continues. Two mechanics are doing the work, and only one of them is about generosity. The first is that the initial bulk training scrape was a one-off event: a lab crawls a large share of the reachable web once, and once it has, its ongoing crawl volume drops toward the rate needed to keep an index fresh. That deflates the numerator permanently. The second is that these products now show and link sources far more often than they did in mid-2025, which inflates the denominator. The first mechanic is largely spent. The second is a product decision that can be reversed in a release. A ratio that fell 95% because a scrape finished is not the same as a ratio that fell because a platform decided to send you traffic, and only the second kind is durable. Plan for the improvement to flatten rather than continue.
Referral share is a separate number, and it is small
The ratio tells you the exchange rate. It says nothing about volume. Those get conflated constantly.
On Cloudflare Radar's referrer data for full-month July 2026, chatgpt.com accounted for 0.913% of crawler referrals against Google's 88.08%. In July 2025 that figure was 0.173%, and in June 2026 it was 0.377%.
The month-on-month jump looks like 2.4x. It is not. TechnologyChecker's own burst analysis shows that 9 to 16 July averaged 2.19% and peaked at 3.035%, while the other 23 days of the month averaged 0.42%. On burst-resistant daily medians the real growth is 1.8x month on month and 2.5x year on year — still fast, roughly half the headline.
| Metric | July 2025 | June 2026 | July 2026 |
|---|---|---|---|
| chatgpt.com share of crawler referrals | 0.173% | 0.377% | 0.913% |
| Google share of crawler referrals | — | — | 88.08% |
| Anthropic crawl-to-refer | 38,744:1 | 3,386:1 | 1,917:1 |
| OpenAI crawl-to-refer | 1,104:1 | 647:1 | 251:1 |
So AI referrals are under 1% of the total and compounding fast, while the exchange rate on those referrals is improving fast. Both things are true, and a policy built on only one of them will be wrong.
The offsetting fact is conversion. Seer Interactive's tracking, cited in the same report, puts ChatGPT-referred traffic converting at 16% against 1.8% for Google organic — while noting that AI traffic is still under 1% of overall organic traffic. A visitor who arrives having already had the product explained to them by a model behaves differently from a visitor who clicked a blue link. We went through that evidence in detail in our analysis of AI search conversion rates and GEO budget.
What Cloudflare's CEO is actually arguing for
Matthew Prince, co-founder and chief executive of Cloudflare, has been blunt about where he thinks this ends. "Clearly it's going to be pay to crawl," he told The Decoder. Announcing the pay-per-crawl mechanism he added: "Now that the majority of traffic on the Internet is non-human, we must go further and act faster so that a sustainable ecosystem can emerge."
Cloudflare followed through: from July 2026 it began defaulting new domains to blocking AI crawlers unless the operator opts in, which pushes AI companies toward paying. Prince's stated design goal is separation: "We hope that our proposed default changes encourage mixed-use crawlers to separate out search from agent use and training."
That last sentence is the one to act on. The useful distinction is not "AI bot versus search bot". It is training crawl versus retrieval crawl versus agent fetch, and those three deserve different answers.
A crawler policy that holds up
Blocking on the basis of a ratio is the wrong instinct, because the ratio is a property of the platform's business model rather than of your site. Decide by crawl purpose instead.
| Crawl purpose | Example agents | Default | Why |
|---|---|---|---|
| Search index that returns clicks | Googlebot, Bingbot | Allow | 4.7:1 and 36.2:1; this is the trade the web has always run |
| Retrieval for a cited answer | OpenAI and Perplexity search agents | Allow | Citations are your only route into an AI answer |
| Live agent fetch on user request | Agent user-agents | Allow | A person asked for your page; this is a visit |
| Bulk training crawl | Training-only user agents | Decide deliberately | No referral path exists; this is a licensing question |
| Unidentified or spoofed | — | Verify then block | Verify by published IP range, not by user-agent string |
Two practical notes.
Blocking retrieval crawlers to protect content is self-defeating in 2026. If your pages are not retrievable, you are not cited, and citations are the only mechanism by which an AI answer sends anyone to you. Our work on earning third-party AI citations makes the same point from the other direction: visibility inside an answer is now the top of the funnel.
Blocking bulk training crawlers is a defensible commercial decision with a real cost — a model that never saw your brand will not mention it unprompted. That is a business call about licensing, not a technical one, and it should be made by whoever owns the content, not by whoever owns the CDN configuration. The mechanics of splitting these categories in Cloudflare are covered in our AI crawler and search agent configuration guide.
Measure your own ratio before you copy anyone's policy
Network-wide averages are a poor guide to your site. A documentation site, a news publisher and a B2B blog have completely different crawl and referral profiles. Compute your own from server logs:
# Your own crawl-to-refer ratio for one platform, from access logs.
# Count HTML responses to their crawler, and HTML views referred by them.
PLATFORM_UA='GPTBot|OAI-SearchBot|ChatGPT-User'
PLATFORM_REF='chatgpt\.com|openai\.com'
CRAWLS=$(awk '$0 ~ /" 200 /' access.log \
| grep -Ec "text/html.*($PLATFORM_UA)")
REFERS=$(awk '$0 ~ /" 200 /' access.log \
| grep -E "text/html" | grep -Ec "$PLATFORM_REF")
echo "crawls=$CRAWLS refers=$REFERS"
[ "$REFERS" -gt 0 ] && echo "ratio=$((CRAWLS / REFERS)):1" \
|| echo "ratio=undefined (no referrals in window)"
Run it over a full calendar month, not a week. TechnologyChecker's report flags that a seven-day sample from 25 May to 1 June 2026 is not comparable with the full-month figures, and week-length windows are exactly how misleading ratios get published. Also remember your own denominator has the same native-app hole Cloudflare's does, so treat your result as a ceiling too.
Pair the log figure with what Search Console now reports; we set out how in tracking GEO visibility in the Search Console AI performance report, and the tooling landscape in GEO measurement and AI citation tracking tools compared.
India-specific considerations
For Indian publishers and D2C brands the calculation tilts further toward allowing retrieval crawlers.
English-language Indian content is comparatively thin in the corpora these models were trained on, particularly for local commercial queries, pricing and regulation. That means retrieval matters more here than in the US: a model answering a question about Indian GST treatment or an Indian pricing band is more likely to be reaching for a live page than reciting something it memorised. Blocking retrieval agents removes you from exactly the queries where you have the least competition.
Second, AI referral volumes in India remain a small share of a large absolute number, so the practical question is not whether to court AI traffic but whether to pay for the infrastructure to serve it. Crawl traffic costs bandwidth. If a training crawler is generating meaningful egress with no referral path, that is a line item worth measuring in ₹ per month before it is worth arguing about in principle.
Third, the Digital Personal Data Protection Act 2023 does not regulate crawling of public pages, but it does regulate what sits on them. Pages exposing personal data to an unauthenticated crawler are a compliance problem regardless of whose bot is reading them, and the crawler policy review is a good moment to check.
What to do this month
Pull your own thirty-day ratio from logs rather than quoting a network average. Separate training crawlers from retrieval and agent fetches in your robots.txt and at the edge, and default retrieval and agent traffic to allow. Re-check the Radar figures each quarter, because a number that moved 95% in twelve months will move again. And treat any ratio you see quoted without a month attached as unusable — the same platform legitimately ranges from 38,744:1 to 1,917:1 depending on which month you picked.
The broader framework for deciding which queries are worth competing for at all sits in our keyword selection framework for AI Overview click survival and the ultimate guide to SEO in 2026.
FAQ
How eCorpIT can help
eCorpIT runs GEO and technical SEO programmes for Indian and global brands, and crawler policy is one of the first things we audit because it silently decides whether a site can be cited at all. Our senior teams pull your real crawl-to-refer figures from logs, split training crawls from retrieval and agent fetches at the edge, and track which AI platforms actually cite you month over month. If you are deciding what to allow and what to charge for, talk to us.
References
- David Belson and Sam Rhea, The crawl before the fall of referrals: understanding AI's impact on content providers, The Cloudflare Blog, 1 July 2025 (updated 15 July 2026).
- Emma Davies, ChatGPT Statistics 2026, TechnologyChecker, updated 4 August 2026, citing Cloudflare Radar
bots/crawlers/summary/crawl_refer_ratiopulled 1 August 2026.
- TechnologyChecker, robots.txt across Cloudflare's network, 2026.
- The Decoder, Cloudflare CEO says the web's future is pay to crawl as bots overtake human traffic, 2025.
- TechCrunch, Cloudflare's new policy pushes AI companies to pay for publishers' content, 1 July 2026.
- SEOmator, GEO Data Report 2026: which AI crawlers and LLM bots take the most and give the least, 2026.
- Digital Applied, AI crawler and bot traffic statistics 2026, 2026.
- SE Ranking, Analysis of top AI search engines: who is catching up to ChatGPT, 2026.
- Goodie, 2026 AI search traffic report, 2026.
- Organik PI, Cloudflare AI Crawl Control: block or charge AI bots, 2026.
- Digiday, The accidental guardian: how Cloudflare's Matthew Prince became publishing's unexpected defender, 2026.
- Fortune, Cloudflare CEO Matthew Prince: Google is abusing its monopoly in search to feed its AI, 13 November 2025.
Last updated: 6 August 2026.