AI Crawler Statistics 2026: Bot Traffic, Blocking, and the Crawl-to-Click Gap

AI crawler traffic is growing, but crawler volume is not AI visibility. Cloudflare’s June 2026 measurement put 52% of crawler requests in the AI-training category (Cloudflare), while its 2025 data showed training dominating AI crawling and sending little or highly uneven referral traffic (Cloudflare). The practical rule is simple: crawl access is not a citation, and a citation is not a click—but neither can happen reliably if the relevant fetcher is blocked.
AI crawler statistics at a glance
These are the findings worth keeping in a dashboard. Each row names the population, date, or method so a number is not mistaken for a universal web average.
| Finding | What was measured | Date, sample, or method |
|---|---|---|
| 52% | Share of crawler requests classified as AI training | Cloudflare crawler telemetry, June 2026 (Cloudflare) |
| 22% | Training share in the earlier comparison point | Cloudflare spring 2025 baseline (Cloudflare) |
| >36% | Share attributed to mixed-use crawlers | Cloudflare taxonomy, June 2026 (Cloudflare) |
| ~80% | AI-crawler activity for training | Cloudflare Radar, prior 12 months ending July 2025 (Cloudflare) |
| 82% | AI-crawler activity for training | Cloudflare Radar, last six months of that 2025 analysis (Cloudflare) |
| 18% | AI-crawler activity for search | Cloudflare Radar, prior 12 months ending July 2025 (Cloudflare) |
| 15% | AI-crawler activity for search | Cloudflare Radar, last six months of that 2025 analysis (Cloudflare) |
| 2% | AI-crawler activity for user actions | Cloudflare Radar, prior 12 months ending July 2025 (Cloudflare) |
| 3% | AI-crawler activity for user actions | Cloudflare Radar, last six months of that 2025 analysis (Cloudflare) |
| 39% | Share of AI-and-search crawler traffic associated with Googlebot | Cloudflare fixed-customer comparison, July 2025 (Cloudflare) |
| 4.7% → 11.7% | GPTBot share of AI crawling traffic | July 2024 to July 2025, Cloudflare Radar (Cloudflare) |
| 6% → ~10% | ClaudeBot share of AI crawling traffic | July 2024 to July 2025, Cloudflare Radar (Cloudflare) |
| 14.1% → 2.4% | Bytespider share of AI crawling traffic | Cloudflare AI-only comparison, July 2024 to July 2025 (Cloudflare) |
| 38,000:1 | Anthropic crawls per referred visitor | July 2025, Cloudflare Radar crawl-to-refer ratio (Cloudflare) |
| 194:1 | Perplexity crawls per referred visitor | July 2025, Cloudflare Radar crawl-to-refer ratio (Cloudflare) |
| 60.0% | Reputable news sites disallowing at least one AI agent | 3,299-site 2025 study; sites with robots.txt analyzed (arXiv) |
| 9.1% | Misinformation sites disallowing at least one AI agent | 672-site 2025 comparison; sites with robots.txt analyzed (arXiv) |
| 15.5 vs 0.77 | Average AI agents named in robots.txt | Same 2025 study: reputable news versus misinformation sites (arXiv) |
What “AI crawler traffic” actually includes
Cloudflare’s newer classification matters because “AI bot” is not one behavior. Training crawlers collect material for future model versions. Search crawlers refresh an index used to ground answers. User-triggered fetchers retrieve a page after a person asks an assistant to inspect it. Mixed-use systems can combine these purposes, so a user-agent label alone may not reveal the commercial consequence of a request.
Cloudflare’s June 2026 report says mixed-use crawlers represented more than 36% of activity, which is precisely why a blanket “allow” or “block” decision can be misleading (Cloudflare, June 2026). A training request can be huge and commercially indirect; a search request can be smaller but determine whether a brand is retrievable and cited today.

Traffic share: training dominates the denominator
Cloudflare’s 2025 Radar analysis found training responsible for about 80% of AI-crawler activity over the preceding year, with search at 18% and user actions at 2%. In its most recent six-month slice, training rose to 82%, search fell to 15%, and user actions reached 3% (Cloudflare, July 2025 analysis).
Those percentages describe AI-crawler activity in Cloudflare’s observed dataset, not all requests on the internet and not the share of human attention. They answer “why are these AI-associated bots crawling?”—not “which brands are visible in answers?”
Crawler mix is changing quickly
Googlebot remained the anchor at 39% of combined AI-and-search crawler traffic in Cloudflare’s fixed-customer comparison for July 2025. Yet AI-specific shares moved materially: GPTBot grew from 4.7% to 11.7% and ClaudeBot from 6% to nearly 10% between July 2024 and July 2025, while Bytespider fell from 14.1% to 2.4% (Cloudflare).
The lesson is not to chase the largest bot. It is to identify which operator supplies the retrieval layer for the answers your buyers use, then check whether that operator can fetch your important pages. Training volume may explain server load or licensing interest; it does not prove that your product is recommended.
Blocking and robots.txt adoption
The 2025 paper Is Misinformation More Open? analyzed 4,079 websites—3,369 reputable news sites and 710 misinformation sites—using 63 AI user agents, seven geographic vantage points, and historical Internet Archive snapshots (study methodology). Among successful sites with robots.txt, 60.0% of reputable sites disallowed at least one AI agent, versus 9.1% of misinformation sites. The averages were 15.5 named agents versus 0.77.
The longitudinal result is more revealing than a single snapshot: reputable-site AI-agent blocking rose from 23% in September 2023 to nearly 60% by May 2025 across six snapshots (arXiv). That is evidence of changing publisher policy, not evidence that blocking improves or harms ranking.
Robots.txt remains a request, not a firewall. Cloudflare explicitly describes compliance as voluntary; its AI Crawl Control documentation adds monitoring for unavailable files, declared violations, and crawler-specific enforcement options (Cloudflare robots.txt docs; AI Crawl Control). Owners can infer intent from directives and observe behavior at the edge, but cannot infer that every unrecognized scraper obeys a token.
The crawl-to-click gap
Cloudflare defines crawl-to-refer ratio as crawler HTML page requests divided by HTML page requests referred back by that platform. In its January–July 2025 table, Anthropic fell from 286,930 crawls per referral in January to 38,066 in July; Perplexity reached 194.8 in July, while Google was 5.4 (Cloudflare Radar analysis).
These are not conversion rates. Native apps may omit a Referer header, and Cloudflare warns that web-only referral counts can overstate the imbalance. A high ratio can mean indexing, training, or an assistant answering without a click. A low ratio can mean genuinely useful referrals—or simply a different measurement boundary.
What site owners can and cannot infer
You can infer that a bot reached a URL, how often it returned, its declared operator, response status, and—when your edge tooling classifies it—an observed purpose. You can compare allowed requests, blocked requests, robots.txt violations, and referrals over the same date range.
You cannot infer that a crawl produced a citation, that a citation produced a recommendation, or that a referral-less answer had no business value. The right measurement stack joins server logs and edge analytics with prompted visibility checks: ask the major assistants buyer questions, record citations and recommendation position, and keep the prompt set stable.
Limitations and update policy
Cloudflare’s figures are observed customer or Radar datasets, not a census; the arXiv study is a curated news/misinformation sample, not every domain. Categories can change as operators disclose new purposes, bot tokens can be spoofed, and crawler-to-referral ratios are sensitive to attribution rules. We will update this page when Cloudflare publishes a comparable new period or the study authors release a materially revised dataset; every figure above retains its original date and method context.
Crawler statistics tell you where access is happening. Run a free AEOeye audit to measure the outcome that traffic logs cannot: whether ChatGPT, Perplexity, Gemini, Google AI, and Claude actually recommend your brand when a buyer asks.
FAQ
What share of crawler requests are for AI training in 2026?+
Cloudflare's classified crawler telemetry shows training is now the largest measured purpose, but the result describes Cloudflare's observed traffic rather than a census of the entire web.
Does more AI crawling mean more website traffic?+
No. Crawl volume and referrals can diverge sharply because training, search indexing, and user-triggered fetching are different activities.
How common is AI-crawler blocking in robots.txt?+
Blocking varies by site type and study sample. Reputable news sites in a recent academic study blocked far more named AI agents than misinformation sites, but the result is not a universal web rate.
Can robots.txt guarantee that an AI crawler cannot access my content?+
No. Robots.txt is a voluntary signal. Technical enforcement requires a control such as a firewall or crawler-control rule.
Sources
- 1.Cloudflare — Content Independence Day, one year on (June 2026 crawler-purpose data)
- 2.Cloudflare — The crawl-to-click gap (Cloudflare Radar, 2025 data)
- 3.Cloudflare — Crawl-to-refer ratios (June 2025 sample)
- 4.Cloudflare AI Crawl Control — Directives
- 5.Cloudflare — Managed robots.txt
- 6.Steinacker-Olsztyn, Gosain & Dao — Is Misinformation More Open? (arXiv 2510.10315)
Is AI recommending you?
Run a free AI visibility audit and find out in under a minute.