AI Crawler Statistics 2026: How Many Websites Block GPTBot, ClaudeBot and CCBot
The story everyone has been telling for two years is that publishers are shutting the door on AI. Our crawl says something narrower: most of the web has not made a decision at all, and most of the sites that appear to have made one were opted in by a vendor. This time we can name the vendor, count its share, and count the sites that decided for themselves.
- Part 1: Robots.txt statistics 2026
- Part 2: AI crawler statistics 2026
- Part 3: llms.txt statistics 2026
- 6,250,526 websites block GPTBot, 6.88% of the 90.8 million with a robots.txt and 59.4% of the sites that name it. Both figures are correct and they are routinely confused.
- 88.3% of the top 1,000 websites allow GPTBot, and so do 88.8% of the top 10,000. The most visited sites block AI crawlers only a little more than the rest of the internet.
- Sites on Cloudflare nameservers block GPTBot at 26.5%. Sites on every other provider block it at 1.3%. 92.7% of all GPTBot blocks with a known DNS provider sit on Cloudflare.
- The eight most blocked AI crawlers are blocked within half a million sites of each other, and 96.0% of the sites carrying Cloudflare's Content-Signals block are on Cloudflare nameservers. The cluster is one managed file.
- 4.3 million websites name GPTBot and allow it, against 6.3 million that name it and block it. Two in five of the sites that wrote a rule for GPTBot wrote an allow.
- 5.9 million websites block OpenAI's training crawler while allowing its search and user-fetch crawlers. Only 13,850 do the reverse.
- 231 thousand websites declare ai-train=yes in a Content-Signals block, an explicit opt-in to AI training. 3.1 million websites with a Content-Signals block do not say ai-train=no at all.
- 54,109 websites block Anthropic's retired anthropic-ai and Claude-Web agents but not ClaudeBot, so the current crawler is allowed. chatgpt.com is one of them.
- Googlebot is blocked by 0.17% of the web and GPTBot by 6.88%. Websites want free traffic and do not want their content used for training with nothing in return, and the gap is likely to widen.
- Companies headquartered in Turkey, India and the Emirates block GPTBot at two to four times the rate of German, Italian and Swiss companies. The United States sits mid-table by headquarters and tops the table by server location only because Cloudflare's edge geolocates there.
How many websites block GPTBot
6.88% of websites with a robots.txt block GPTBot: 6,250,526 of the 90.8 million sites that serve one. Among the 10.5 million sites that name GPTBot at all, 59.4% block it.
Across the whole web, all 158.2 million live websites, about 4% block GPTBot, because a site with no robots.txt allows every crawler.
Those three percentages describe one fact and sit tens of points apart. Nearly every published AI blocking figure is one of them without saying which. In this report "the web" means the 90.8 million sites that publish a robots.txt, and every share is a share of them unless it says otherwise.
The most blocked AI crawlers
| Agent | Based in | What it does | Sites blocking it | Share of the web | Of sites naming it | Sites naming it |
|---|---|---|---|---|---|---|
| USA | Model training | 6,250,526 | 6.88% | 59.4% | 10,519,282 | |
| USA | Open archive | 6,172,210 | 6.80% | 61.6% | 10,013,330 | |
| USA | Training, Alexa | 6,124,858 | 6.75% | 62.7% | 9,773,093 | |
| China | Model training | 6,070,738 | 6.69% | 62.4% | 9,727,955 | |
| USA | Model training | 5,965,470 | 6.57% | 59.3% | 10,067,682 | |
| USA | Gemini opt-out | 5,877,087 | 6.47% | 59.4% | 9,900,472 | |
| USA | Apple AI opt-out | 5,768,884 | 6.35% | 61.1% | 9,446,162 | |
| USA | Model training | 5,737,436 | 6.32% | 60.9% | 9,424,643 | |
| USA | User fetch | 391,822 | 0.43% | 30.1% | 1,300,175 | |
| USA | Retired for ClaudeBot | 385,950 | 0.43% | 9.3% | 4,134,418 | |
| USA | Answer engine | 340,333 | 0.37% | 25.5% | 1,335,284 | |
| Israel | Data resale | 314,353 | 0.35% | 8.8% | 3,560,098 | |
| USA | Answer engine | 299,531 | 0.33% | 8.4% | 3,554,799 | |
| Canada | Model training | 258,862 | 0.29% | 6.9% | 3,736,802 | |
| USA | Image datasets | 257,347 | 0.28% | 67.3% | 382,632 | |
| USA | Retired for ClaudeBot | 257,156 | 0.28% | 32.4% | 794,603 | |
| USA | Model training | 235,393 | 0.26% | 6.4% | 3,704,497 | |
| UK | Company data | 199,924 | 0.22% | 6.0% | 3,330,274 | |
| USA | Web extraction | 176,390 | 0.19% | 38.9% | 453,362 | |
| USA | ChatGPT search | 130,680 | 0.14% | 14.2% | 917,959 | |
| Image datasets | 39,217 | 0.04% | 1.2% | 3,199,644 |
Embed this figure
The eight most blocked AI crawlers sit within half a million sites of each other. Of the 6.2 million sites that block CCBot, Common Crawl's fetcher, 5.7 million also block GPTBot, ClaudeBot, Bytespider and Google-Extended in the same file. 5.9 million sites block eight or more named crawlers.
Sites do not arrive at identical eight-item block lists on their own. Something wrote the list for them, and the next section names it.
| Agent | Sites allowing it | Sites blocking it | Allow | Block | Sites naming it |
|---|---|---|---|---|---|
| 125,285 | 257,347 | 32.7% | 67.3% | 382,632 | |
| 3,648,235 | 6,124,858 | 37.3% | 62.7% | 9,773,093 | |
| 3,657,217 | 6,070,738 | 37.6% | 62.4% | 9,727,955 | |
| 3,841,120 | 6,172,210 | 38.4% | 61.6% | 10,013,330 | |
| 3,677,278 | 5,768,884 | 38.9% | 61.1% | 9,446,162 | |
| 3,687,207 | 5,737,436 | 39.1% | 60.9% | 9,424,643 | |
| 4,268,756 | 6,250,526 | 40.6% | 59.4% | 10,519,282 | |
| 4,023,385 | 5,877,087 | 40.6% | 59.4% | 9,900,472 | |
| 4,102,212 | 5,965,470 | 40.7% | 59.3% | 10,067,682 | |
| 276,972 | 176,390 | 61.1% | 38.9% | 453,362 | |
| 537,447 | 257,156 | 67.6% | 32.4% | 794,603 | |
| 908,353 | 391,822 | 69.9% | 30.1% | 1,300,175 | |
| 994,951 | 340,333 | 74.5% | 25.5% | 1,335,284 | |
| 787,279 | 130,680 | 85.8% | 14.2% | 917,959 | |
| 3,748,468 | 385,950 | 90.7% | 9.3% | 4,134,418 | |
| 3,245,745 | 314,353 | 91.2% | 8.8% | 3,560,098 | |
| 3,255,268 | 299,531 | 91.6% | 8.4% | 3,554,799 | |
| 3,477,940 | 258,862 | 93.1% | 6.9% | 3,736,802 | |
| 3,469,104 | 235,393 | 93.6% | 6.4% | 3,704,497 | |
| 3,130,350 | 199,924 | 94.0% | 6.0% | 3,330,274 | |
| 3,160,427 | 39,217 | 98.8% | 1.2% | 3,199,644 |
Embed this figure
Once a website has written a rule for a crawler, the decision splits two ways, and the split is the chart above. Amazonbot, Bytespider and CCBot are blocked by more than three in five of the sites that name them, and the chart keeps to crawlers named by at least a million websites, so a crawler almost nobody has heard of cannot top it on a handful of mentions. GPTBot and ClaudeBot sit at 59.4% and 59.3%.
The answering crawlers are much lower: ChatGPT-User at 30.1%, PerplexityBot at 25.5% and OAI-SearchBot at 14.2%. Sites that write a rule for a training crawler mostly block it. Sites that write a rule for a crawler that answers a person's question mostly allow it.
Two entries in the chart are not crawlers. Google-Extended and Applebot-Extended fetch nothing. Googlebot and Applebot do the crawling, and the token tells them whether the pages may be used to train Gemini or Apple Intelligence. Google's documentation says Google-Extended "does not impact a site's inclusion in Google Search".
So 5.9 million sites refuse Google-Extended and 150.9 thousand refuse Googlebot. Same company, same file, 38.9 times the take-up. Saying no to Google-Extended costs a publisher nothing in search. Saying no to Googlebot costs the search traffic itself.
That is the whole bargain in one comparison. Googlebot is blocked by 0.17% of the web and GPTBot by 6.88%: websites want free traffic, and they do not want their content used for training with nothing offered in return. Until a training crawler sends visitors, expect that gap to widen.
Cloudflare wrote most of the blocks
Embed this figure
Among the 34.7 million sites whose nameservers we can identify, the ones on Cloudflare block GPTBot at 26.5% and the ones everywhere else at 1.3%. 92.7% of all GPTBot blocks sit on Cloudflare, and 96.0% of the sites carrying Cloudflare's Content-Signals block are on Cloudflare nameservers.
Take Cloudflare's customers out and the sentence for the rest of the internet becomes: about one website in eighty blocks GPTBot.
| Nameservers | Block GPTBot |
|---|---|
| On Cloudflare | 26.5% |
| Everywhere else | 1.3% |
Embed this figure
| DNS provider | Sites | Block GPTBot | Carry Content-Signals | Closed to unlisted crawlers | Publish llms.txt | Declare a sitemap | Sites blocking GPTBot |
|---|---|---|---|---|---|---|---|
| Cloudflare | 13,194,969 | 26.5% | 40.2% | 4.8% | 10.2% | 55.3% | 3,501,945 |
| AWS Route 53 | 751,025 | 4.9% | 1.5% | 6.5% | 17.5% | 72.0% | 36,650 |
| Namecheap | 1,357,950 | 4.0% | 3.8% | 7.5% | 20.8% | 72.0% | 54,590 |
| WordPress | 1,051,992 | 3.1% | 0.2% | 13.0% | 3.7% | 84.0% | 32,507 |
| OVH | 769,128 | 1.3% | 0.7% | 1.9% | 10.2% | 80.5% | 9,999 |
| GoDaddy | 9,994,787 | 1.0% | 1.1% | 2.5% | 36.2% | 89.5% | 97,949 |
| NS1 | 615,200 | 0.9% | 0.7% | 1.1% | 4.8% | 93.2% | 5,660 |
| 2,117,314 | 0.9% | 0.9% | 7.2% | 39.2% | 84.0% | 18,632 | |
| Squarespace | 877,352 | 0.3% | 0.3% | 0.6% | 2.8% | 98.8% | 2,369 |
| Wix | 3,798,175 | 0.2% | 0.2% | 0.1% | 90.7% | 99.4% | 7,596 |
Embed this figure
Cloudflare ships a managed robots.txt that writes the Content Signals Policy preamble and a block list of AI crawlers into the file, and it can be turned on for a whole account in one setting. GoDaddy, Wix, Squarespace and Google, the other providers that write robots.txt for customers by the million, all sit at or under one percent. Their templates do other things, and none of them blocks an AI crawler by default.
A default that nobody turns off is still a preference, and Cloudflare shipping it changed the internet's posture faster than any publisher campaign did. But the headline number is usually cited as several million separate editorial decisions, and it is not that. The share of the web that blocks AI crawlers will move with the customer counts of a few infrastructure companies, not with publisher opinion, and the next big jump will come from whichever platform switches a block on by default.
What the Cloudflare block actually declares
| Declared in a Content-Signals block | Sites | Share of files |
|---|---|---|
| Any Content-Signals block | 8,621,685 | 9.5% |
| ai-train=no | 5,502,794 | 6.1% |
| ai-train=yes | 231,031 | 0.3% |
| search=yes | 5,731,476 | 6.3% |
| search=no | 1,891 | 0.0% |
| ai-input=yes | 292,478 | 0.3% |
| ai-input=no | 13,464 | 0.0% |
| A block that does not say ai-train=no | 3,119,169 | 3.4% |
Embed this figure
The Content-Signals block is usually described as "blocking AI". 5.5 million websites say ai-train=no, but 3.1 million websites with a Content-Signals block do not say it at all, and 231 thousand say ai-train=yes: an explicit, machine-readable opt-in to model training that no published count has reported.
6.3% of the web says search=yes, 5.7 million of them, and only 1,891 say search=no.
Deliberate decisions are rare, and they lean towards allowing
Strip out the managed block and a much smaller population appears. It is the one worth quoting.
4.7% of the web, 4,268,756 websites, names GPTBot in a rule that allows it, against the 6,250,526 that name it and block it. A crawler that is named in the file follows its own rules and ignores User-agent: *, so a GPTBot rule that does not deny the whole site is a permission somebody typed. Two in five of the sites that wrote a rule for GPTBot wrote an allow.
ClaudeBot is allowed by name on 4.1 million sites and Google-Extended on 4 million.
Sites that distinguish one AI crawler from another are rarer still. 226,505 block CCBot while allowing GPTBot and ClaudeBot, a position against a research archive but not the model builders. 354,849 do the opposite. Together they are 0.2% and 0.4% of the web. Roughly one website in two hundred has drawn a line between AI companies. Everyone else blocks all of them or none.
Where a website blocks exactly one AI crawler, the choice is telling. 737.7 thousand websites do it: GPTBot on 88,455, Bytespider on 43,091, Meta-ExternalAgent on 41,931, Amazonbot on 40,132, ClaudeBot on 9,976 and PerplexityBot on 2,531. Bytespider's place near the top is notable because ByteDance's crawler is the one most widely reported to ignore robots.txt, so a block aimed only at it is a statement rather than a control.
Training blocked, answering allowed: OpenAI's three crawlers
OpenAI runs three crawlers: GPTBot for training, OAI-SearchBot for ChatGPT search, and ChatGPT-User, which fetches a page because a person asked a question. OpenAI asks sites to treat them separately, and the internet does.
Embed this figure
| Position on OpenAI's three agents | Sites |
|---|---|
| Refuse GPTBot, admit OAI-SearchBot and ChatGPT-User | 5,861,912 |
| Refuse all three | 132,874 |
| Admit GPTBot, refuse the answering agents | 13,850 |
Embed this figure
6.5% of the web, 5,861,912 websites, blocks GPTBot while allowing both OAI-SearchBot and ChatGPT-User. 132,874 block all three. Only 13,850 allow the trainer and block the answering crawlers.
Most of the first group is the Cloudflare default, whose managed list blocks training crawlers and leaves search and user-fetch crawlers alone. The shape holds outside it too: sites are not shutting down, they are separating training from answering. Websites are pricing the trade one crawler at a time. A fetch that brings a reader is worth allowing; a copy taken for training is not, and the line between the two is where the next robots.txt rules will be written.
Do the biggest websites block AI crawlers
| Popularity tier | Block GPTBot | Sites blocking | Sites in tier |
|---|---|---|---|
| Top 1,000 | 11.7% | 97 | 830 |
| Top 10,000 | 11.2% | 1,052 | 9,371 |
| Top 100,000 | 10.8% | 9,780 | 90,404 |
| Top 1 million | 9.3% | 88,806 | 953,471 |
| Entire internet | 6.9% | 6,250,526 | 90,796,807 |
Embed this figure
The gradient is real and small. 11.7% of the top 1,000 websites block GPTBot against 6.9% across the internet. Put the other way round: 88.3% of the top 1,000 websites allow GPTBot, and so do 88.8% of the top 10,000.
| Agent | Top 100k | Top 1M | Whole internet | Concentration |
|---|---|---|---|---|
| 10.7% | 9.3% | 6.9% | 1.6x | |
| 10.1% | 9.1% | 6.8% | 1.5x | |
| 8.4% | 8.3% | 6.8% | 1.3x | |
| 9.5% | 8.6% | 6.7% | 1.4x | |
| 9.2% | 8.4% | 6.6% | 1.4x | |
| 8.9% | 8.2% | 6.5% | 1.4x | |
| 8.2% | 7.7% | 6.4% | 1.3x | |
| 8.4% | 7.8% | 6.3% | 1.3x | |
| 2.4% | 1.4% | 0.4% | 5.5x | |
| 2.4% | 1.3% | 0.4% | 5.5x | |
| 2.1% | 1.1% | 0.4% | 5.7x | |
| 1.8% | 0.9% | 0.4% | 5.0x | |
| 1.7% | 0.9% | 0.3% | 5.1x | |
| 1.8% | 1.0% | 0.3% | 6.3x | |
| 1.4% | 0.9% | 0.3% | 5.1x | |
| 1.8% | 1.1% | 0.3% | 6.5x | |
| 1.8% | 0.9% | 0.3% | 6.9x | |
| 0.4% | 0.2% | 0.2% | 2.0x | |
| 1.8% | 0.7% | 0.2% | 9.2x | |
| 1.1% | 0.4% | 0.1% | 8.1x |
Embed this figure
What large sites single out is not training. PerplexityBot is blocked by 0.4% of the internet and 2.1% of the top hundred thousand sites. ChatGPT-User goes from 0.4% to 2.4%. The training crawlers barely move by comparison.
Those two are answer engines. They fetch a page because a person asked and show the answer instead of the link. A training crawler takes a copy once. An answer engine takes the reader every time, and the sites with the most readers to lose noticed first. The next fight over robots.txt is about answer engines, not training, and the largest publishers are already fighting it.
Where the blocking companies are
| Company headquarters | Sites | Block GPTBot | Block CCBot | Sites blocking GPTBot |
|---|---|---|---|---|
| Turkey | 96,973 | 8.6% | 8.6% | 8,311 |
| India | 349,820 | 6.5% | 6.6% | 22,843 |
| UAE | 74,331 | 6.4% | 6.5% | 4,787 |
| Brazil | 209,078 | 5.8% | 5.8% | 12,106 |
| UK | 518,931 | 4.7% | 4.6% | 24,597 |
| Canada | 244,212 | 4.7% | 4.6% | 11,502 |
| Australia | 252,705 | 4.4% | 4.5% | 11,043 |
| USA | 1,832,035 | 4.2% | 4.2% | 77,129 |
| Mexico | 70,621 | 4.2% | 4.2% | 2,952 |
| Poland | 70,757 | 4.0% | 4.0% | 2,852 |
| Spain | 184,218 | 3.3% | 3.3% | 6,024 |
| Netherlands | 229,272 | 3.1% | 3.0% | 7,153 |
| France | 282,009 | 2.8% | 2.8% | 7,896 |
| Switzerland | 96,187 | 2.7% | 2.5% | 2,626 |
| Italy | 188,369 | 2.5% | 2.6% | 4,747 |
| Germany | 307,056 | 2.2% | 2.2% | 6,663 |
Embed this figure
Matching sites to company records gives a cut the popularity ranking cannot: who owns the site. Turkish, Indian and Emirati companies block GPTBot at 8.6%, 6.5% and 6.4%. German, Italian and Swiss companies sit at 2.2%, 2.5% and 2.7%. American companies are in the middle at 4.2%.
Whatever drives the decision, it is not the volume of national debate about AI and copyright.
A cut by server location gives a different answer, and it is wrong in an instructive way. By the country of the IP address that answered, the United States blocks GPTBot at 10.8% and every other country sits between half a percent and four. Cloudflare's edge answers for sites all over the world and its addresses geolocate to the United States, so the server-country cut measures the Cloudflare default a second time. The headquarters chart uses the company's own stated location and is immune to that.
Bigger companies block more
| Company headcount | GPTBot | CCBot | Cloudflare signals | Googlebot |
|---|---|---|---|---|
| 1 person | 4.0% | 4.0% | 4.7% | 0.0% |
| 2 to 10 | 4.7% | 4.7% | 5.8% | 0.0% |
| 11 to 50 | 4.9% | 4.8% | 6.0% | 0.0% |
| 51 to 200 | 5.0% | 4.9% | 5.8% | 0.0% |
| 201 to 500 | 5.2% | 5.1% | 5.8% | 0.0% |
| 501 to 1,000 | 5.4% | 5.3% | 5.9% | 0.0% |
| 1,001 to 5,000 | 5.2% | 5.1% | 5.5% | 0.0% |
| 5,001 to 10,000 | 5.3% | 5.0% | 6.0% | 0.1% |
| over 10,000 | 6.5% | 6.3% | 6.9% | 0.0% |
Embed this figure
Every step up in company size raises the share that blocks GPTBot, from 4.0% at one-person companies to 6.5% above ten thousand staff. The Content-Signals block follows the same curve, from 4.7% to 6.9%, which is enterprise plans carrying enterprise defaults.
Googlebot blocking stays under 0.03% at every size, and that flat line is the control. Whatever rises with company size, it is not a general appetite for shutting crawlers out. A large company has a legal team, and a legal team writes rules about training data, not about search.
54,109 websites block a crawler that no longer exists
Anthropic retired the anthropic-ai and Claude-Web user agents in favour of ClaudeBot, and its documentation names only the new one. The old names are still everywhere.
0.43% of the web, 385,950 websites, blocks anthropic-ai, and 257,156 block Claude-Web. Most also block ClaudeBot, so the stale lines are harmless. But 54,109 websites block only the retired agents, 14.0% of everyone who wrote the old name down. Those publishers believe they have opted out of Anthropic's crawler. ClaudeBot is allowed on every one of them.
OpenAI's own chatgpt.com blocks Anthropic's two dead agents and does not name the live one, so ClaudeBot follows the wildcard group and is allowed everything the wildcard allows. So are 461 of the top hundred thousand websites.
A robots.txt is written once and rarely revisited, and crawler names change without notice. Any block list copied from an article published before mid-2024 carries agents that have since been renamed or retired. As long as AI companies rename crawlers faster than websites re-read their robots.txt, a growing share of the web will be blocking ghosts.
Frequently asked questions
How many websites block GPTBot?
6.88% of websites with a robots.txt block GPTBot: 6,250,526 of the 90,796,807 sites serving one, and 59.4% of the 10,519,282 sites that mention GPTBot at all.
Which AI crawler is blocked most?
By count, GPTBot at 6.3 million sites, with CCBot, Amazonbot, Bytespider and ClaudeBot within half a million of it. By block rate among sites that name it, Amazonbot at 62.7%, then Bytespider and CCBot, all above 60%.
Why do sites on Cloudflare block AI crawlers so much more?
Cloudflare offers a managed robots.txt that writes a Content-Signals block and an AI crawler block list into the file, switched on for an account in one setting. Sites on Cloudflare nameservers block GPTBot at 26.5%; sites on every other provider block it at 1.3%.
How many websites block Common Crawl?
6.80% of websites with a robots.txt block CCBot, 6,172,210 sites. Only 226,505 block CCBot while allowing GPTBot and ClaudeBot, so almost nobody singles Common Crawl out.
Do the biggest websites block AI crawlers?
More than average, but not by much. 11.7% of the top 1,000 websites block GPTBot against 6.9% across the internet, and the sites they single out are the answer engines rather than the training crawlers.
Does blocking a crawler in robots.txt actually stop it?
Only crawlers that choose to obey it. It also only works if the user agent string in the file matches the one the crawler sends, which is why 54,109 websites block retired Anthropic agents while ClaudeBot is allowed in.
Do any websites opt in to AI training?
Yes. 231 thousand websites carry ai-train=yes in a Content-Signals block, and 4.3 million name GPTBot in a rule that allows it.
Methodology and sources
All figures come from the StackScan September 2026 crawl, which covers the entire internet: we requested /robots.txt, /llms.txt, /ads.txt and /sitemap.xml from 158,187,361 live websites in the September 2026 crawl. 90,796,807 returned a robots.txt we could parse, and that count is the denominator wherever a percentage is given "of sites with a robots.txt".
Rules are read per user-agent group, following RFC 9309. A crawler counts as blocked when its own group contains a bare Disallow: /; a crawler whose own group restricts some paths is not blocked, whatever the wildcard group says. Naming a crawler is recorded separately from blocking it, because Content-Signals blocks list many crawlers they then allow, and treating presence as a block overstates blocking by several million sites. Named sites and popularity tiers use the same whole-site rule.
The DNS provider is read from each site's nameservers at crawl time, one provider per domain, and is known for 34.7 million of the sites with a robots.txt. Server country comes from the IP address that answered. Popularity tiers come from the StackScan domain ranking. Company headquarters and size come from StackScan's company records, matched to sites by website address. That match covers a minority of sites and leans towards businesses, so the company charts describe businesses rather than the whole web.
Two limits. The 67.4 million live sites with no robots.txt are excluded from every percentage, and an absent file allows everything, so the true permissive share of the internet is higher than the figures here. And this is one three-week window, not a trend; a fixed panel is frozen so the next run can be compared with this one.
Every chart has a table view, and every figure can be cited with a link back to it. Figures are free to reuse with attribution. The crawlers above are documented by their operators: OpenAI, Anthropic, Google and Common Crawl. The Content Signals Policy is published at contentsignals.org.