GEO Data Report 2026: Which AI Crawlers & LLM Bots Take the Most and Give the Least?

Mistral crawls 3,389 pages for every referral it sends back — now the most extractive AI bot on the web, ahead of Anthropic (2,237:1) and OpenAI (217:1). I cross-checked Cloudflare Radar's global network against SEOmator's own 500+ site panel to rebuild the crawl-to-refer ratio for July 2026.

Published
GEO Data Report 2026: Which AI Crawlers & LLM Bots Take the Most and Give the Least?

Mistral's crawler now fetches 3,389 pages for every referral its platform sends back to a website owner — making it, as of July 2026, the most extractive AI bot on the open web. Anthropic's ClaudeBot, which held that title earlier this year, has fallen to 2,237:1; OpenAI's GPTBot sits at 217:1; DuckDuckGo runs at near-parity, 2.5:1. Those figures come from Cloudflare Radar's rolling 28-day window ending July 21, 2026 — and I verified every one of them against a completely independent dataset: SEOmator's own measurement network of 500+ sites and 50M+ monthly users. The two networks agree almost line-for-line, which is what makes the metric at the center of this report worth trusting: the crawl-to-refer ratio — the number of pages an AI crawler or LLM bot fetches divided by the referrals its parent platform (ChatGPT, Claude, Perplexity, Mistral, Copilot) sends back.

Data sources: Cloudflare Radar — Bot & Crawler Analytics (rolling 28-day window ending July 21, 2026), cross-checked against SEOmator's own crawling and analytics network (500+ sites monitored, 50M+ users tracked monthly, January–July 2026).

GEO data report 2026 showing which AI crawlers and LLM bots have the highest crawl-to-refer ratios

What Is the Crawl-to-Refer Ratio and Why Does It Matter for GEO and SEO?

The crawl-to-refer ratio measures how many pages an AI crawler or LLM bot crawls from your website for every referral visit it sends back. A ratio of 100:1 means the bot crawled 100 of your pages before its parent platform directed a single user to your site.

This metric matters because AI crawlers — GPTBot, ClaudeBot, PerplexityBot, and others — consume your server resources (bandwidth, compute, and crawl budget) while indexing your content for their LLM models. The referral side measures whether that investment pays off through actual traffic returned to your site.

For SEO and GEO professionals, this ratio functions as a return-on-investment metric for crawler access. Every page crawled by an LLM bot is a page that could have been crawled by Googlebot instead, directly affecting your site's crawl budget allocation. And the load is no longer a rounding error: across the 500+ sites SEOmator monitors, AI-related bots (crawlers, assistants, and AI-search fetchers combined) grew from roughly 18% of all bot traffic in January to about 34% by July 2026 — nearly one in three bot requests, on a business-web panel that blocks aggressively.

How Do the Top AI Crawlers and LLM Bots Compare on Crawl-to-Refer Ratios?

The gap between the best and worst crawl-to-refer ratios still spans three orders of magnitude, but the ranking at the top has changed since the start of the year. Mistral — a rounding error on this chart in the spring — has surged to the most extractive position, while Anthropic and OpenAI, once the runaway leaders, have pulled their ratios down by an order of magnitude.

OperatorCrawl-to-Refer Ratio (Jul 2026)What This Means
Mistral3,389:1Crawls 3,389 pages per 1 referral — the new most-extractive operator
Anthropic (ClaudeBot)2,237:1Down sharply from 5-figure ratios earlier in 2026
Perplexity (PerplexityBot)225:1Rising — one of the few operators getting more extractive
OpenAI (GPTBot)217:1Collapsed from four figures as ChatGPT Search matured
Microsoft (Copilot)35:1Includes Copilot and Bing AI features
Yandex26:1Russian search with growing AI features
Baidu12:1Chinese search with established referral patterns
ByteDance11:1TikTok's parent — referrals rising through H1
Google (Gemini / AI Overviews)4.6:1Traditional search drives high referral volume
DuckDuckGo2.5:1Near-parity — the most efficient ratio of all operators

Mistral and Anthropic sit at the extractive end for the same structural reason: both are primarily model-training operations whose consumer surfaces send back little or no clickable referral traffic. Google and DuckDuckGo sit at the efficient end because their core function is to send users to websites — every search result click is a referral. I've been tracking these operators in server logs for the better part of a year, and the disparity is still striking when you see it at the individual-site level.

Two independent networks, one story

Cloudflare Radar measures Cloudflare's global network. SEOmator's Agent Analytics measures a different population entirely — 500+ mostly B2B sites, 50M+ users a month — computing the same ratio from its own crawl logs and its own referral sessions. When two networks with no shared plumbing land on the same numbers, the signal is real. Here is July 2026 side by side:

OperatorCloudflare Radar (global)SEOmator panel (B2B-weighted)
Mistral3,389:16,021:1
Anthropic2,237:12,363:1
Perplexity225:1263:1
OpenAI217:1179:1
Microsoft35:135:1
Google4.6:14.0:1
DuckDuckGo2.5:12.7:1

The agreement on Anthropic (2,237 vs 2,363), Microsoft (35 vs 35), and Google (4.6 vs 4.0) is close enough to rule out either dataset being an artifact. Both also independently flag Mistral as the new extraction leader and Perplexity as a quiet outlier moving the wrong way. Understanding this ratio is now a core part of any serious AI search optimization strategy.

Why Did the Extraction Gap Collapse in 2026?

Google's structural advantage in AI search: traditional search sends users to sites while training crawlers do not

Earlier in 2026, Anthropic's ratio was a five-figure number and OpenAI's was well over a thousand. The reason both fell so far is not that they crawl less — it's that their platforms finally started sending traffic back. As ChatGPT Search and Claude's cited-source surfaces matured, referrals appeared on the denominator of the ratio for the first time, and the number dropped fast.

SEOmator's panel captured the descent month by month, because it sees both halves of the equation at once:

OperatorJanMarMayJul*
Anthropic56,969:111,870:111,934:12,363:1
OpenAI1,264:11,416:11,225:1179:1
Perplexity116:1114:1175:1263:1
Mistral22:139:140:16,021:1
Google5.1:15.5:15.0:14.0:1

*July reflects a partial month (July 1–21) on the SEOmator panel.

The Anthropic line is the clearest signal in the dataset: an operator that took roughly 57,000 pages per referral in January was down to the low thousands by July as Claude's user-facing surfaces began returning clicks. OpenAI followed the same path from four figures to the mid-100s. This is the referral side of the bargain reasserting itself — and it validates a claim I've made to clients all year: extraction ratios are not fixed, and they move the moment a platform ships a product that links out.

Why Is Mistral Suddenly the Most Extractive Operator?

Mistral is the mirror image of the Anthropic story. For the first five months of 2026 its crawler was barely visible and its ratio sat in the low double digits. Then, from June, its crawl volume spiked across both networks while its platform sent back virtually no referral traffic — pushing its ratio to 3,389:1 on Cloudflare's global network and past 6,000:1 on SEOmator's B2B-weighted panel.

That combination — a sharp increase in crawling with no reciprocal referral mechanism — is exactly the profile that made ClaudeBot the poster child for extraction a year ago. If you are auditing AI crawler access this quarter, Mistral's user agents (MistralAI-User, and Mistral's training crawler) are the new names to watch in your logs. The lesson from Anthropic is that the ratio can improve once a consumer product ships; the question for publishers is whether to subsidize the training phase in the meantime.

How Much of Your Crawl Load Is Actually AI — and What Are the Bots Fetching?

Before deciding what to block, it helps to see the composition of the traffic. On Cloudflare's network, humans and AI bots now split the request stream like this:

Client TypeShare of Traffic (Jul 2026)
Non-AI Bot48.6%
Human42.4%
AI Bot5.7%
Mixed Purpose3.3%

The 5.7% "AI bot" figure understates the pressure on content sites specifically. SEOmator's panel — weighted toward the docs, blogs, and changelogs that make prime training corpus — shows AI bots hitting B2B SaaS sites at 38% of all bot traffic (July 2026 snapshot), with more than half of that crawling declared as training-purpose. E-commerce sites see relatively less training crawling but proportionally more search/answer retrieval, because product queries pull live pages.

What are all these bots actually pulling down? This breakdown wasn't in the original edition of this report, and it's one of the more useful additions. By declared crawl purpose across AI bots:

Crawl PurposeShare of AI-Bot Requests
Training46.5%
Mixed purpose39.3%
Search / answer retrieval11.1%
User-triggered action2.5%
Undeclared0.6%

Training is now the single largest declared purpose at 46.5% — and SEOmator's panel shows the same drift, from 39% in January to roughly 46% by July, while search/answer crawling climbed from 8% to 12%. The practical reading: AI bots increasingly crawl your site to answer live user questions, not only to train — which is precisely the traffic you want to be citable for.

And by content type, AI crawlers are overwhelmingly after your prose, not your assets:

Content TypeShare of AI-Bot Requests
HTML72.9%
JSON6.9%
JavaScript5.6%
Images5.4%
Plain Text5.1%
XML / CSS / other4.2%

Nearly three-quarters of AI crawler requests target HTML. Your written content — articles, documentation, product descriptions — is the corpus. That is both the risk (it's being taken) and the opportunity (it's what gets cited).

Which AI Crawlers and LLM Bots Actually Crawl the Most Pages?

Looking at total user-agent share across all crawler traffic (not just AI), Googlebot still leads — but the bot in second place is no longer a search engine. It's Claude-User, Anthropic's on-demand agent that fetches pages when a Claude user asks a question.

User AgentShare of All Crawler Traffic (Jul 2026)Category
Googlebot12.97%Search Engine
Claude-User (Anthropic)10.10%AI Assistant
Meta-ExternalAgent (Meta AI)6.35%AI Crawler / LLM Bot
GPTBot (ChatGPT)5.38%AI Crawler / LLM Bot
Bingbot4.93%Search Engine
Google AdsBot3.68%Advertising
facebookexternalhit3.55%Social Preview
Applebot3.38%Search / AI
Amazonbot3.32%AI Crawler / LLM Bot

Within the AI-crawler population specifically, the Anthropic-over-OpenAI reversal is even starker. ClaudeBot now takes an 18.2% share of AI-and-adjacent crawler requests — nearly double GPTBot's 9.8% — and Anthropic's crawlers combined (ClaudeBot plus Claude-SearchBot) are the largest AI-specific operator on the web. SEOmator's panel recorded the same crossover: ClaudeBot overtook GPTBot in June and holds the top AI-crawler slot through July. This is a fundamental shift from a year ago, when OpenAI was the default name in every AI-crawler conversation. Our earlier research on AI SEO statistics anticipated this crossover, though Anthropic's specific surge outran the forecast.

Where Does Referral Traffic Actually Come From?

Despite all that crawling, referral traffic is still overwhelmingly Google. But the notable line here is ChatGPT — and it moved faster than the original edition of this report predicted.

ReferrerShare of All Referral Traffic (Jul 2026)
Google88.15%
Bing3.10%
TikTok2.43%
Yandex1.77%
DuckDuckGo1.26%
ChatGPT1.05%
Baidu (mobile)0.89%
Bing China (cn.bing.com)0.61%
Baidu (desktop)0.46%

ChatGPT's share of referrals has crossed 1% — up roughly 5× from the 0.20% this report recorded in the spring. The earlier edition projected chatgpt.com reaching 1% "by late 2026"; it got there by mid-July, alongside DuckDuckGo and Baidu as a measurable traffic source.

SEOmator's analytics panel tells the referral story from the other side — as a share of sessions rather than referrer hits — and the trend rhymes:

AI EngineShare of All Sessions (Jul 2026, SEOmator panel)
All AI engines1.19%
— ChatGPT0.65%
— Perplexity0.20%
— Gemini0.15%
— Claude0.10%

Across the sites SEOmator tracks, AI engines sent about 1.2% of all sessions in July — nearly double January's 0.6%, and the fastest-growing channel on the panel. ChatGPT drives roughly 55% of it. The vertical cut is the strategic part: B2B SaaS sites see about 1.6% of sessions from AI answers versus 0.7% for e-commerce — AI surfaces send researchers, not shoppers. Pair that with the falling extraction ratios above and the picture is coherent: as crawlers got less extractive, the referrals finally showed up.

Which Industries Are Crawled the Most by AI Crawlers and LLM Bots?

Retail and infrastructure content absorb a disproportionate share of crawler attention relative to their footprint on the web.

IndustryShare of Crawler Traffic (Jul 2026)
Shopping & General Merchandise26.4%
Internet & Telecom20.6%
Computer & Electronics18.8%
News, Media & Publications9.2%
Gambling6.7%
Business & Industry3.6%
Professional Services2.7%
Finance2.5%

Shopping sites absorb over a quarter of all crawler traffic — and, as the extraction ratios showed, receive some of the least referral return per crawl. That double penalty makes e-commerce the biggest single subsidizer of LLM model training. SEOmator's B2B-weighted panel sees the same pattern from a different starting mix: shopping is about 20% of its customers but roughly 21% of all crawler hits, so per-site, retail punches well above its weight. When I run technical SEO audits for SaaS clients, industry-specific crawl patterns are always the first thing I check — a finance SaaS and a consumer-electronics retailer get fundamentally different returns from the same set of bots. I documented related AI bot traffic patterns by country in an earlier analysis, where geography drove comparable variation.

Is the Web Actually Ready for the AI Agents Crawling It?

Here is a gap that reframes the whole "should I block?" question. SEOmator ran its audit engine's AI-readiness checks across roughly 109,000 top domains in July 2026; Cloudflare's independent scan of ~106,000 domains found nearly identical numbers. Both tell the same story:

SignalShare of Domains (Jul 2026)
robots.txt with AI-crawler rules~80–81%
XML sitemap~69–70%
Serves markdown to agents~7%
MCP server card (agent-native)~0.3%
Web Bot Auth<0.1%

Roughly four in five top sites now write explicit AI-crawler rules — yet fewer than 7% serve clean markdown to agents, and agent-native protocols (MCP, Web Bot Auth) are effectively at zero. The web is writing rules for AI agents it isn't actually equipped to serve well. If you're deciding how to handle AI crawlers, that's the real state of the art: most sites are still fighting the last war (allow/deny in robots.txt) instead of the next one (serving agents structured, efficient content).

How Many Websites Are Blocking AI Crawlers and LLM Bots?

Website operators are responding to unfavorable ratios by naming AI crawlers in their robots.txt files — and the two names at the top of every list are the two biggest extractors. SEOmator's crawler parsed 4,257 top-domain robots.txt files on July 20, 2026:

AgentFiles Naming ItOperator
GPTBot773OpenAI
ClaudeBot629Anthropic
Google-Extended578Google (AI opt-out)
CCBot568Common Crawl
Bytespider507ByteDance
meta-externalagent495Meta
Amazonbot477Amazon
PerplexityBot470Perplexity
Applebot-Extended420Apple (AI opt-out)
ChatGPT-User375OpenAI

Cloudflare's independent robots.txt scan produces the same top two — GPTBot (738 directives) and ClaudeBot (652) — confirming that AI opt-out is now the dominant reason sites write crawler rules at all. Technology and business domains lead the blocking, exactly as they did in the spring:

Domain CategoryDomains With AI-Crawler Rules (Cloudflare scan)
Technology934
Business786
Trackers / Analytics317
E-commerce300
Search Engines263
Content Servers202

Notably, blocking is rising even as the ratios improve. On SEOmator's panel, the crawler rejection rate (4xx responses) climbed from about 35% in January to 37% by mid-year — more than one in three crawler requests now turned away — and the share actually served (2xx) has slipped below 50%. If you're unsure how your own robots.txt is configured for LLM bots, review our guide on what LLMs.txt is and how to generate it as part of a broader AI crawler access strategy.

Should You Block AI Crawlers and LLM Bots Based on This Data?

The decision depends on your industry, traffic goals, and long-term GEO strategy. Based on the July 2026 data, the operators sort into four groups — and the membership has shifted since the start of the year.

Block with confidence: Mistral (3,389:1) and Meta-ExternalAgent. Mistral is now the most extractive crawler on the web with no meaningful referral product, and Meta-ExternalAgent has never had one. These provide near-zero return and consume resources exclusively for the operator's benefit.

Evaluate carefully: ClaudeBot (2,237:1) and GPTBot (217:1). Both still train models on your content, but both have improved by an order of magnitude as Claude and ChatGPT Search began returning traffic. Blocking them now also means forfeiting representation in the two fastest-growing AI answer surfaces — a real AI search optimization cost that didn't exist when their ratios were five figures.

Allow strategically: Microsoft Copilot (35:1). Moderate ratio, growing referral traffic through Bing and Copilot surfaces, and prominent source citation. Watch PerplexityBot (225:1): it's one of the few operators getting more extractive in 2026, so it's moved from "allow" toward "evaluate."

Keep unblocked: Google (4.6:1), DuckDuckGo (2.5:1), and traditional search crawlers. These deliver measurable referral traffic that justifies their crawl volume.

One caveat before you write any robots.txt rule: referrals are not the only return on being crawled. A bot can shape how an assistant describes your brand without ever sending a click, so blocking is not a free action. And for the operators you do decide to allow, access is not the only lever — you can generate an llms.txt file to tell those bots which of your pages actually matter, rather than leaving them to infer it from a full crawl.

What Does This Mean for the Future of GEO, AI Search, and Publisher Relations?

The crawl-to-refer data from the first half of 2026 revises the story the web told itself in 2025. The fear was that LLM platforms would crawl endlessly and never send traffic back. What actually happened is more nuanced: the biggest extractors improved dramatically the moment they shipped consumer products that link out — while a new operator (Mistral) took over the extraction lead in their place. Three trends point to where this is heading:

1. Extraction ratios are elastic, not fixed. Anthropic went from ~57,000:1 to ~2,300:1 in six months, and OpenAI from four figures to the mid-100s, purely because referrals appeared. Any operator that launches a cited-answer product can follow the same curve — which means today's worst offender is not necessarily tomorrow's.

2. ChatGPT referrals crossed 1% ahead of schedule. The spring edition of this report projected that milestone for late 2026; it arrived in July. As AI answer engines mature, GEO optimization for them stops being speculative and starts showing up in analytics.

3. The readiness gap is the next battleground. With ~80% of sites writing AI rules but only ~7% serving markdown to agents, the winners of the next phase won't be the sites that block hardest — they'll be the ones that serve AI agents efficient, structured, citable content. Blocking is defense; being citable is offense.

Website owners who monitor these ratios quarterly and adjust their AI crawler access based on measurable returns will outperform those who either block everything or allow everything. That is the foundation of a data-driven GEO strategy.

How to Monitor Your Site's Crawl-to-Refer Ratio

While Cloudflare Radar and SEOmator provide network-wide aggregates, individual site owners can approximate their own crawl-to-refer ratios using server logs and analytics. Here's the process I follow for clients:

  1. Server log analysis: Count requests from known AI crawler and LLM bot user agents (GPTBot, ClaudeBot, Claude-User, MistralAI-User, Meta-ExternalAgent, PerplexityBot) over a 30-day period
  2. Referral tracking: In Google Analytics 4 or your analytics platform, filter referral traffic from chatgpt.com, perplexity.ai, gemini.google.com, and other LLM-powered surfaces
  3. Calculate the ratio: Divide total AI crawler requests by total AI-referred visits, per operator
  4. Compare against benchmarks: Use the July 2026 ratios in this report as baselines, and re-check quarterly — as the Anthropic and Mistral swings show, the numbers move fast

If you don't have server-log access, an AI crawler access report covers the crawl half of the equation — it shows which LLM bots your robots.txt currently admits or blocks, which is the input you can actually change.

SEO tools like SEOmator can help automate the technical side — analyzing your robots.txt configuration, monitoring AI crawler access patterns, and identifying which LLM bots are consuming your crawl budget without delivering proportional value.

Methodology

This report draws on two independent datasets, deliberately cross-checked against each other.

Cloudflare Radar monitors traffic across Cloudflare's global network, a significant share of all internet traffic, making its bot and crawler data broadly representative of web-wide patterns. All Cloudflare figures use a rolling 28-day window ending July 21, 2026.

SEOmator's figures come from its own crawling and analytics network: 500+ sites monitored by Agent Analytics (50M+ users tracked monthly), a 100K+ site audit engine, and a prompt tracker running 1M+ keywords and prompts monthly, over January–July 2026. Because this panel is weighted toward B2B SaaS (~60% of customers), its shares describe how the business web behaves, not the consumer internet at large — which is exactly why it's a useful second opinion on Cloudflare's broader network.

Data dimensions used:

  • Crawl-to-refer ratio: Pages crawled per referral sent, by operator (both networks)
  • Referral source share: Referral traffic by source domain (Cloudflare) and AI-engine session share (SEOmator)
  • Client type / bot category: Traffic classification (Human, Non-AI Bot, AI Bot, Mixed)
  • User agent + crawl purpose + content type: Crawler identification and what AI bots fetch
  • Vertical / industry: Traffic segmentation by website category
  • Agent readiness + robots.txt analysis: AI-rule and agent-protocol adoption across ~106K–109K domains

Date range: Cloudflare — 28 days ending July 21, 2026. SEOmator — January 1 – July 21, 2026 (July is a partial month, July 1–21).

Limitations: Network-wide ratios measure aggregate behavior; individual-site ratios vary with content type, domain authority, and traffic patterns. Referral attribution may undercount AI-driven visits that arrive through intermediate pages or omit a referrer header. SEOmator's panel is B2B-weighted by design and reports shares across the sites it tracks, not the internet as a whole.

Cloudflare data sourced from Cloudflare Radar Bot & Crawler Analytics (28 days ending July 21, 2026). Source: SEOmator crawling and analytics data, January–July 2026 (500+ sites monitored / 50M+ users tracked monthly / 1M+ prompts & keywords tracked monthly). Last updated: July 21, 2026.

Explore more stories