Six months ago, when I first analyzed this data, Googlebot commanded 38.7% of all AI-related crawler traffic and the world split neatly into two camps. That map has already been redrawn. Across the 28 days ending July 23, 2026, Cloudflare Radar shows Googlebot down to 24.4%, ClaudeBot risen to #2 at 17.9%, and the tidy geographic story fractured into three patterns instead of two -- with ByteDance's Bytespider now leading an entire major market. Just as striking: only 46% of AI-crawler requests are answered with a 200 OK. More than a quarter are met with a 403 Forbidden or a 429 Too Many Requests. The web has started fighting back, and it is doing so at very different intensities depending on where your audience sits.
I re-ran the full analysis across 11 countries -- by user agent, crawl purpose, content type, HTTP response, and robots.txt directives -- to see exactly how the AI-crawler landscape looks in mid-2026 and how it shifted since February.
What Does the Global AI Bot Landscape Look Like in Mid-2026?
Googlebot still leads all AI-related crawler traffic globally, but its lead has collapsed to 24.4% -- down nearly 14 percentage points from the 38.7% I measured in February. ClaudeBot has surged into second at 17.9%, and if you group Anthropic's two crawlers together (ClaudeBot plus Claude-SearchBot at 3.5%), Anthropic's combined 21.4% share now rivals Google's entire footprint. This is according to Cloudflare Radar (ai/bots/summary/user_agent), covering websites protected by Cloudflare's network of 330+ cities in 125+ countries.
| AI Bot | Global Share (%) | Operator | Primary Purpose |
|---|
| Googlebot | 24.4% | Google | Search indexing + AI training (mixed) |
| ClaudeBot | 17.9% | Anthropic | Model training |
| Meta-ExternalAgent | 11.8% | Meta | AI training |
| GPTBot | 9.7% | OpenAI | Model training |
| Bingbot | 8.8% | Microsoft | Search indexing + AI (mixed) |
| Bytespider | 6.3% | ByteDance | AI training |
| Applebot | 6.2% | Apple | Search + AI features |
| Amazonbot | 6.0% | Amazon | AI training |
| Claude-SearchBot | 3.5% | Anthropic | Claude web search |
The top four crawlers now account for 63.8% of identified AI bot requests -- a less concentrated market than February's 74.3%, because the runners-up have grown while Google shrank. The single biggest story is Anthropic: ClaudeBot went from the #4 spot (11.4%) to a clear #2 (17.9%) in six months, and Anthropic launched a dedicated search crawler (Claude-SearchBot) that already pulls a 3.5% global share on its own.
How Does AI Bot Traffic Differ by Country?
In February this section described a clean binary -- a "Googlebot Belt" versus a "ChatGPT Corridor." That framing no longer survives contact with the data. Brazil and Japan have moved into the Googlebot camp, Australia's ChatGPT dominance has more than halved, and a genuinely new pattern has appeared: Bytespider now leads India outright. There are three patterns now, not two.
| Country | #1 Bot | #1 Share | #2 Bot | #2 Share | Training % | User Action % |
|---|
| Brazil | Googlebot | 50.8% | ChatGPT-User | 25.5% | 12.0% | 27.3% |
| Germany | Googlebot | 42.8% | Bytespider | 10.7% | 27.4% | 12.0% |
| Canada | Googlebot | 42.5% | Bytespider | 23.5% | 35.3% | 2.0% |
| France | Googlebot | 36.6% | Meta-ExternalAgent | 10.5% | 33.2% | 3.4% |
| United Kingdom | Googlebot | 30.4% | ChatGPT-User | 9.3% | 31.2% | 11.9% |
| United States | Googlebot | 25.6% | ClaudeBot | 19.8% | 44.4% | 1.0% |
| Japan | Googlebot | 23.6% | ChatGPT-User | 23.5% | 24.4% | 29.1% |
| Netherlands | Googlebot | 23.3% | Meta-ExternalAgent | 12.2% | 38.5% | 5.8% |
| South Korea | ChatGPT-User | 27.3% | Googlebot | 25.5% | 14.6% | 30.8% |
| Australia | ChatGPT-User | 36.5% | Googlebot | 21.1% | 22.4% | 40.7% |
| India | Bytespider | 29.6% | Googlebot | 21.1% | 41.6% | 16.5% |
Pattern 1: Googlebot Majority Markets
Googlebot holds the #1 position in eight of the eleven countries I analyzed. Brazil is now the most Googlebot-heavy market at 50.8% -- a complete reversal from February, when ChatGPT-User led Brazil at 57.6%. Germany (42.8%) and Canada (42.5%) round out the top of the belt. The United States, once solidly Googlebot territory at 40.3%, has dropped to 25.6% as ClaudeBot climbed to a very close second at 19.8% -- meaning in the US, Anthropic's crawler is now nearly as active as Google's.
Pattern 2: ChatGPT-User Markets
Australia and South Korea remain the two countries where ChatGPT-User -- the bot that visits pages when real users ask ChatGPT questions -- is the dominant crawler. But the intensity has cooled sharply. Australia's ChatGPT-User share fell from a staggering 75.8% in February to 36.5% today, and its "User Action" crawl purpose dropped from 76.0% to 40.7%. It is still the most user-driven AI-crawler market on Earth, but no longer a near-monopoly. Japan is now a photo-finish: Googlebot 23.6% versus ChatGPT-User 23.5%, with the highest User Action share of any Googlebot-led market (29.1%).
Pattern 3: The Bytespider Markets (New)
The clearest new signal in the data is ByteDance's Bytespider leading India at 29.6% -- ahead of Googlebot's 21.1%. Bytespider is also the #2 crawler in Canada (23.5%) and Germany (10.7%). This is a training-first pattern: India's crawl activity is 41.6% "Training," the highest of any country I looked at outside the US. Where ChatGPT-User traffic signals end-user demand for AI answers, a Bytespider surge signals bulk data collection for model training -- a different pressure on your infrastructure, and one that returns no referral traffic at all.
Why Does the Map Keep Redrawing Itself?
The February two-camp model broke because the AI-crawler market is still forming. Three forces are visibly at work in the country-level data:
- New entrants scale faster than incumbents. Anthropic's ClaudeBot and Claude-SearchBot, Meta-ExternalAgent, and Bytespider all grew their share while Googlebot's proportional footprint shrank -- not because Google crawls less, but because everyone else crawls more. A ranking built on relative share reshuffles every month during a land grab.
- User-driven crawling follows consumer habit, not headquarters. Markets like Australia, South Korea, and Japan show high "User Action" percentages because ChatGPT adoption is high and competing local AI services are few. When someone in Sydney or Seoul asks ChatGPT a question, the bot fetches a page in real time -- so consumer behavior, not the location of AI labs, drives that traffic.
- Training crawlers chase language and volume. English-heavy markets (US, Canada) and large content markets (India, Germany) draw the training-first crawlers -- GPTBot, ClaudeBot, Meta-ExternalAgent, and Bytespider -- which is why those countries skew "Training" and "Mixed Purpose" rather than "User Action."
What Content Are AI Bots Actually Fetching?
A breakdown Cloudflare Radar exposes but few analyses use is content type (ai/bots/summary/content_type) -- the MIME category of what AI crawlers request. Globally, AI bots overwhelmingly fetch rendered pages, but the non-HTML slice is where the interesting behavior hides.
| Content Type | Share of AI Bot Requests |
|---|
| HTML | 72.8% |
| JSON | 7.0% |
| JavaScript | 5.6% |
| Images | 5.3% |
| Plain Text | 5.2% |
| XML | 1.7% |
| CSS | 1.4% |
Nearly three-quarters of AI-crawler requests are for HTML -- the readable page. But 7.0% is JSON and 5.2% is plain text, and together that 12%+ tells you AI crawlers are increasingly pulling structured and raw data, not just prose: API responses, JSON-LD payloads, sitemaps, and the plain-text files (llms.txt, robots.txt, feeds) that agents parse directly. If you want AI systems to consume your facts cleanly rather than guessing from rendered HTML, a machine-readable surface matters -- our free llms.txt generator builds one that these crawlers can read without executing your JavaScript.
How Often Do Websites Actually Serve AI Bots?
Here is the finding that reframes the entire "should I block AI bots" debate. Cloudflare Radar's response-status breakdown (bots/crawlers/summary/response_status) shows the HTTP status codes crawlers actually receive -- and websites are already saying no, at scale.
| HTTP Response | Share of AI Bot Requests | Meaning |
|---|
| 200 OK | 46.0% | Served successfully |
| 403 Forbidden | 20.7% | Access denied (blocked) |
| 301 Moved | 8.1% | Permanent redirect |
| 404 Not Found | 7.8% | Missing page |
| 429 Too Many Requests | 6.0% | Rate-limited (throttled) |
| 302 Found | 4.6% | Temporary redirect |
| 204 No Content | 1.6% | Empty success |
| 503 Unavailable | 1.3% | Server overloaded / blocking |
| 304 Not Modified | 1.2% | Cached, unchanged |
Only 46% of AI-crawler requests get a clean 200. A combined 26.7% are actively refused or throttled (20.7% 403 + 6.0% 429), and another 1.3% hit a 503. In other words, roughly one in four AI-crawler requests is already being turned away by the sites it targets -- through WAF rules, bot-management products, or origin rate limits. The "AI is scraping everything unchecked" narrative is out of date; a large and growing share of the web is refusing these bots by default. That also means AI answer engines are increasingly working from partial crawls -- if your competitors block and you don't, your pages are the ones getting cited.
Which Crawlers Are Websites Trying to Govern?
If response codes show how sites react in real time, robots.txt shows their stated intent. Cloudflare Radar parses robots.txt files across its domain population and ranks the user agents those files name most often (robots_txt/top/user_agents/directive). The list is a near-perfect ranking of AI crawlers -- a direct signal of which bots site owners are consciously writing rules for.
| User Agent | Domains Naming It in robots.txt | Operator |
|---|
| GPTBot | 743 | OpenAI |
| ClaudeBot | 656 | Anthropic |
| Google-Extended | 626 | Google (AI training opt-out) |
| CCBot | 603 | Common Crawl |
| Bytespider | 523 | ByteDance |
| meta-externalagent | 465 | Meta |
| PerplexityBot | 448 | Perplexity |
| Amazonbot | 446 | Amazon |
| Applebot-Extended | 442 | Apple (AI training opt-out) |
| Googlebot | 416 | Google |
These are counts of domains that name each agent in a directive (a single-day snapshot), not traffic shares -- a rule can allow or restrict, and the ranking counts both.
The striking part is who tops the list. GPTBot (743) and ClaudeBot (656) are named in more robots.txt files than Googlebot itself (416) -- the classic search crawler that has been around for two decades. Site owners are writing rules specifically for the AI-training crawlers, and the presence of Google-Extended (626) and Applebot-Extended (442) -- the opt-out tokens Google and Apple created specifically so publishers can refuse AI training while keeping search indexing -- confirms this is a deliberate governance effort, not incidental config.
One caveat that makes this actionable: a robots.txt directive only works if it is syntactically valid and targets the right token. GPTBot and ClaudeBot honor robots.txt; Bytespider's compliance is inconsistent. Whichever you decide to allow or block, validate your robots.txt file afterwards -- a malformed rule fails open, and you will believe you are blocking a crawler that is walking right through.
Which Industries Do AI Bots Target Most?
According to Cloudflare Radar (bots/crawlers/summary/vertical), shopping and e-commerce remain the single most-crawled vertical, but its dominance has eased as AI crawlers spread across other verticals.
| Industry Vertical | Share of AI Bot Traffic |
|---|
| Shopping and General Merchandise | 26.2% |
| Internet and Telecom | 20.6% |
| Computer and Electronics | 18.9% |
| News, Media, and Publications | 9.1% |
| Gambling | 6.7% |
| Business and Industry | 3.6% |
| Professional Services | 2.7% |
| Finance | 2.5% |
| Games | 2.3% |
Retail's 26.2% (down from 31.2% in February) still leads, which tracks with how heavily AI models are used for product research and comparison shopping -- when a user asks an assistant "what's the best running shoe under $150," the bot needs current product pages and reviews to answer. But note the rise of Internet/Telecom (20.6%) and Computer/Electronics (18.9%): technical documentation and product specs are prime training and retrieval material, and those verticals are being crawled almost as hard as retail now.
What Are AI Bots Actually Doing With the Content They Crawl?
Cloudflare Radar classifies AI bot activity into purpose categories (ai/bots/summary/crawl_purpose). The global split has tilted decisively toward training since February:
- Training (46.2%): Dedicated training crawlers -- GPTBot, ClaudeBot, Meta-ExternalAgent, Bytespider -- explicitly collecting data to build and improve AI models. This is now the largest category, overtaking Mixed Purpose.
- Mixed Purpose (39.4%): Crawlers like Googlebot and Bingbot that index for search and collect training data at the same time, without cleanly separating the two.
- Search (11.2%): AI-powered search bots (OAI-SearchBot, Claude-SearchBot) that crawl pages to generate search answers -- up from 6.9% in February as AI search products matured.
- User Action (2.5%): Bots like ChatGPT-User that fetch pages in real time when users ask questions. Small globally, but concentrated in specific markets (40.7% of Australia's crawl activity).
The country-level split remains sharp. In the United States, 86.4% of AI bot activity is Training or Mixed Purpose -- US websites are primarily raw material for model building. In Australia, 40.7% of activity is User Action -- the majority of the crawl is triggered by real people asking AI for answers.
How Has AI Bot Traffic Trended?
AI-crawler request volume has climbed steadily through the first half of 2026. The chart below tracks Cloudflare Radar's AI-bot traffic index over the trailing 24 weeks:
Source: Cloudflare Radar ai/bots/timeseries (normalized index, peak = 1.0). The index rose from 0.78 in early February to 0.95 by late July 2026 -- roughly a 21% increase in AI-crawler volume over six months.
Zooming into the last month, the composition kept shifting even as the total grew. Comparing the current 28-day window (June 25 – July 23, 2026) against the prior 28 days (May 28 – June 25):
| AI Bot | Previous 28d Share | Current 28d Share | Change |
|---|
| Googlebot | 25.9% | 24.4% | -1.5 pp |
| ClaudeBot | 18.2% | 17.9% | -0.3 pp |
| Meta-ExternalAgent | 10.5% | 11.8% | +1.3 pp |
| GPTBot | 9.4% | 9.7% | +0.3 pp |
| Bingbot | 7.8% | 8.8% | +1.0 pp |
| Bytespider | 8.0% | 6.3% | -1.7 pp |
| Amazonbot | 5.3% | 6.0% | +0.7 pp |
| Claude-SearchBot | 3.1% | 3.5% | +0.4 pp |
Meta-ExternalAgent (+1.3 pp) and Bingbot (+1.0 pp) were the month's biggest gainers. Bytespider dropped 1.7 pp globally in the last month even as it dominates India -- a reminder that a crawler can be fading in aggregate while surging in a single market. For the sharper six-month view, the headline is Anthropic: ClaudeBot roughly doubled its slice of the AI-crawler market between February and July.
What Should Website Owners Do About AI Bot Traffic?
The blocking decision ultimately turns on a single number: how many pages a bot takes for every visitor it sends back. Our crawl-to-refer ratio analysis ranks every major operator on exactly that -- a spread of four orders of magnitude, from DuckDuckGo's near-parity to Anthropic's lopsided extreme. With that ratio in hand, tailor your response to your market:
If your audience is in the US, Canada, or Western Europe:
- Expect Googlebot and, increasingly, ClaudeBot to be your primary AI crawlers. In the US the two are now nearly tied. Because Google bundles search indexing with AI training under one bot, a robots.txt block on the AI portion is not cleanly possible -- your leverage there is limited.
- Watch Meta-ExternalAgent and Bytespider. Both grew as second-place crawlers in European markets; if your logs show Bytespider climbing, know its robots.txt compliance is unreliable and you may need a WAF rule rather than a polite directive.
- GPTBot and ClaudeBot honor robots.txt. Blocking both removes roughly a quarter of dedicated AI-training traffic. Whichever way you go, validate the file -- a malformed directive fails open.
If your audience is in Australia, South Korea, or Japan:
- ChatGPT-User is your primary concern and opportunity. In Australia it drives 40.7% of AI-crawler activity through real user questions. Blocking it means your content will not appear when someone in your market asks ChatGPT -- a growing referral source, not a training cost.
- Training crawlers are a smaller share here, so the "AI is stealing my content" problem is less about model training and more about whether you want real-time AI-search visibility.
- Optimize for AI citation. Since most AI-bot traffic in your market is user-triggered, structuring content with clear answers, statistics, and authoritative sources increases the odds of being quoted. A GEO audit for AI search engines shows how well your pages currently hold up on those signals.
If your audience is in India or other training-heavy markets:
- Bytespider may be your single largest AI crawler. It is training-first and returns no referral traffic, so the cost-benefit skews toward controlling it -- but plan for enforcement (rate limits, bot management) rather than robots.txt alone, given its inconsistent compliance.
For everyone -- expect to be refused, and design for it:
- With only 46% of AI-bot requests globally receiving a 200 and 26.7% getting a 403 or 429, the crawl your competitors are surviving may be one you are silently blocking -- or vice versa. Audit your own logs for AI user agents and the status codes you return them; an unintended block can quietly erase you from AI answers.
- Serve a clean machine-readable surface (structured data, sitemaps, llms.txt) so the crawls you do allow extract your facts accurately instead of guessing from rendered HTML.
How I Analyzed This Data
This analysis uses Cloudflare Radar AI-bot data pulled directly from the Radar REST API. The primary window covers the 28 days from June 25 to July 23, 2026; the month-over-month comparison uses the preceding 28-day control window (May 28 – June 25, 2026), and the trend chart uses the trailing 24 weeks. February 2026 figures cited for comparison come from my original analysis of the same dataset.
Endpoints queried: ai/bots/summary/user_agent, ai/bots/summary/crawl_purpose, and ai/bots/summary/content_type for global and per-country breakdowns; bots/crawlers/summary/vertical and bots/crawlers/summary/response_status for industry and HTTP-status data; robots_txt/top/user_agents/directive for the robots.txt governance snapshot; and ai/bots/timeseries for the trend index. Country breakdowns cover 11 markets: United States, United Kingdom, Germany, France, Netherlands, Canada, Japan, Australia, South Korea, India, and Brazil.
All percentages represent share of identified AI bot requests, not share of total web traffic, except where noted: the robots.txt figures are domain counts (a snapshot), and the trend line is a normalized index (peak = 1.0), not an absolute request count. The location filter corresponds to the billing country of the Cloudflare customer whose site received the traffic, aggregated across Cloudflare's global network of 330+ cities in 125+ countries. Values shift daily; re-run the queries before quoting them.
Data source: Cloudflare Radar -- radar/ai/bots/*, radar/bots/crawlers/*, and radar/robots_txt/* endpoints (radar.cloudflare.com), 28 days ending July 23, 2026. Last updated: July 24, 2026.