llms.txt Statistics 2026: What 17.9 Million Websites Publish
llms.txt has real adoption and the adoption figure is quoted everywhere. Almost nobody has looked at who wrote them or what is in them. These are the numbers, counted across the entire internet: how many sites have one, which platforms generated them, what they say, and how many contradict the site's own robots.txt.
- Part 1: Robots.txt statistics 2026
- Part 2: AI crawler statistics 2026
- Part 3: llms.txt statistics 2026
- 11.4% of the web has an llms.txt: 17,956,155 of the internet's 158,187,361 live websites. That is far above every published estimate, and the rest of this list is why.
- Wix generated the llms.txt of 3.5 million websites and GoDaddy of 3.6 million, two fifths of the total between them. Another 2.2 million websites, 12.1%, serve Shopify's storefront template.
- 19.9% of websites with an llms.txt, 3,570,999 of them, serve one that is empty or describes a site that is not open: launching soon, store unavailable, or an HTML page served under the name.
- 723,109 websites use a Disallow-Training directive that appears in no specification. 722,605 of them are one identical line written by GoDaddy's site builder.
- Three WordPress SEO plugins wrote the llms.txt of 1.8 million websites, 10.2% of the total. 2.3 million websites serve one that uses the word "generated" about itself.
- The most linked destinations across every llms.txt are shop.app, dev.wix.com and shopify.com. The single most linked destination across the web's llms.txt is a platform's own property, not any site's content.
- 1.1 million websites serve an llms.txt over 200 bytes that contains no links at all. The median one has 5 links.
- 548.4 thousand websites publish an llms.txt and block every unlisted crawler in robots.txt. 407 thousand publish one and block GPTBot.
- 26.3% of one-person companies publish an llms.txt against 15.0% of companies above ten thousand staff. A standard pitched at publishers is being adopted by hosting defaults on the smallest sites.
How many websites have an llms.txt
llms.txt was proposed in 2024 as a Markdown file at the root of a site that tells a language model what the site is and where its useful content lives. 11.4% of the web has one. We fetched /llms.txt from the entire internet, 158,187,361 live websites, and 17,956,155 returned one.
That is many times higher than the published estimates, which sample large publishers or count GitHub repositories. The gap is explained by who wrote them.
Who wrote them: Wix, GoDaddy and Shopify
Embed this figure
| DNS provider | Sites with an llms.txt | Files | Sites on the provider |
|---|---|---|---|
| Wix | 58.4% | 3,452,053 | 5,910,420 |
| GoDaddy | 24.4% | 3,642,299 | 14,949,650 |
| 14.5% | 835,005 | 5,748,774 | |
| AWS Route 53 | 9.4% | 135,568 | 1,445,201 |
| Namecheap | 7.4% | 287,318 | 3,877,044 |
| Cloudflare | 7.1% | 1,391,646 | 19,607,873 |
| OVH | 5.9% | 80,600 | 1,358,987 |
| WordPress | 3.4% | 39,178 | 1,155,337 |
| NS1 | 3.0% | 20,940 | 689,322 |
| Squarespace | 1.8% | 34,566 | 1,922,950 |
Embed this figure
Sites on Wix nameservers serve an llms.txt at 58.4%. Wix generates the file for its site builder customers, and among Wix sites that also publish a robots.txt, the closer proxy for a published site, the rate is 90.7%.
GoDaddy's site builder generates one on 24.4% of the sites on GoDaddy nameservers. Squarespace, at 1.8%, and WordPress.com, at 3.4%, made the opposite decision, and their customers have almost none.
| Who wrote the file | Files | Share of llms.txt |
|---|---|---|
| Mentions Wix | 4,328,723 | 24.1% |
| GoDaddy (DNS) | 3,642,299 | 20.3% |
| Wix (DNS) | 3,452,053 | 19.2% |
| Says "generated" | 2,266,801 | 12.6% |
| Mentions Shopify | 2,169,403 | 12.1% |
| Cloudflare (DNS) | 1,391,646 | 7.8% |
| All in One SEO | 1,118,229 | 6.2% |
| Google (DNS) | 835,005 | 4.7% |
| Yoast SEO | 523,198 | 2.9% |
| Namecheap (DNS) | 287,318 | 1.6% |
| Rank Math SEO | 197,580 | 1.1% |
Embed this figure
Wix wrote the llms.txt of 3,452,053 websites and GoDaddy of 3,642,299, 19.2% and 20.3% of the total. 2.2 million websites, 12.1%, serve Shopify's generated version, which addresses AI shopping agents rather than search engines.
The three WordPress SEO plugins, All in One SEO, Yoast and Rank Math, wrote 1.8 million between them, 10.2%. 2.3 million websites serve one that describes itself with the word "generated".
This is the WordPress finding from part one arriving in a younger file. The most widely deployed crawl instruction on the internet is a default nobody chose, and the most widely deployed AI instruction now is too. Any llms.txt adoption figure measures which platforms shipped a generator, not what publishers want, and the number will jump every time a large platform switches one on.
What websites put in it
| What the file is | Sites | Share of llms.txt |
|---|---|---|
| Empty | 1,723,099 | 9.6% |
| Under 200 bytes | 3,591,513 | 20.0% |
| Headings and links, 200 bytes or more | 11,521,675 | 64.2% |
| Says the site is launching soon | 1,224,112 | 6.8% |
| Says the store is unavailable | 429,921 | 2.4% |
| HTML or a script, not text | 193,867 | 1.1% |
Embed this figure
19.9% of websites with an llms.txt, 3,570,999 of them, serve one that is empty or belongs to a site that is not open: nothing at all, a parked domain announcing a launch, a closed storefront, or an HTML page served under the wrong name.
The 429.9 thousand "Store Unavailable" versions are Shopify's placeholder for a password-protected or suspended shop, faithfully describing a store nobody can enter. GoDaddy serves an llms.txt of zero bytes on 198.5 thousand websites.
The counterweight matters. 11,521,675 websites, 64.2%, serve one that is genuinely structured: Markdown headings, links, two hundred bytes or more. A clear majority is a real attempt at the format, even where a platform made the attempt. The adoption figure everyone quotes counts the placeholders and the templates too.
723,109 websites use a directive that does not exist
4.4% of websites with an llms.txt, 791,861 of them, have a User-agent: line in it. That is robots.txt syntax in a file specified as Markdown, and it does nothing. 723,109 of them go further and carry a Disallow-Training: directive. It appears in no specification and is honoured by no AI operator.
Those owners believe they have opted out of model training. They have not, in the same way the 54,109 sites in part two blocking Anthropic's retired user agents have not. This group is thirteen times larger.
It is also not a snippet that spread from blog post to blog post. 722,605 of those 723,109 websites begin the file with the identical line, User-agent: * Allow: / Disallow-Training: / Sitemap: /sitemap.xml.
By nameservers, 369.8 thousand of those sites are on GoDaddy, and most of the rest are on the two networks that host GoDaddy's website builder. The directive that does not exist was written by GoDaddy's builder once and shipped to every customer who published a site. The demand behind it is real: website owners want a permissions file for AI and reached for the nearest thing. Nobody on the AI side has given them one.
| What the file contains | Files | Share of llms.txt |
|---|---|---|
| A User-agent line | 791,861 | 4.4% |
| Disallow-Training | 723,109 | 4.0% |
| of which one identical template line | 722,605 | 4.0% |
| Commercial-use: allowed | 28,612 | 0.2% |
| Lorem ipsum | 203,360 | 1.1% |
| The word "please" | 2,541,214 | 14.2% |
| Chinese characters | 623,946 | 3.5% |
| Cyrillic | 186,914 | 1.0% |
| A link to llms-full.txt | 138,911 | 0.8% |
| A link to llmstxt.org | 12,243 | 0.1% |
| Over 1 MB | 4,426 | 0.0% |
Embed this figure
A second invented set, Commercial-use: allowed and Research-use: allowed, appears on 28.6 thousand websites, another generator's template. 203.4 thousand websites serve lorem ipsum in it, the placeholder text of a template nobody finished, handed to every language model that asks. Only 12.2 thousand link to the specification they follow.
Where the links point
A link is the whole point of the format. Across every llms.txt of two hundred bytes or more we counted 161,904,107 of them, a mean of 12.8 per website and a median of 5. Most are a very short map, and 1,100,303 websites serve one over two hundred bytes with no links whatsoever.
| Most linked destination | Links pointing at it |
|---|---|
| shop.app | 6,365,583 |
| dev.wix.com | 4,312,752 |
| ucp.dev | 2,122,452 |
| www.shopify.com | 2,121,703 |
| www.facebook.com | 405,857 |
| www.instagram.com | 334,102 |
| beacons.ai | 271,480 |
| docs.github.com | 246,302 |
| github.com | 179,417 |
| www.rediff.com | 155,790 |
| support.myclickfunnels.com | 140,760 |
| www.linkedin.com | 124,691 |
Embed this figure
Strip out every link to the site's own domain and rank what is left. The top of the chart is not a publisher's content. Shopify's shop.app and shopify.com and Wix's developer documentation at dev.wix.com account for tens of millions of links inside llms.txt that are supposed to describe somebody else's website.
The single most linked destination across 18 million websites is a platform's own property. A file meant to guide language models to a site's best pages is, at scale, guiding them to the platform's checkout.
Adoption by company size
| Company headcount | Publish an llms.txt | Sites publishing | Sites in band |
|---|---|---|---|
| 1 person | 26.3% | 216,123 | 822,072 |
| 2 to 10 | 25.5% | 1,077,306 | 4,223,072 |
| 11 to 50 | 22.6% | 505,146 | 2,237,138 |
| 51 to 200 | 20.2% | 143,326 | 708,833 |
| 201 to 500 | 18.1% | 37,151 | 205,027 |
| 501 to 1,000 | 17.2% | 12,753 | 74,103 |
| 1,001 to 5,000 | 15.3% | 8,563 | 55,896 |
| 5,001 to 10,000 | 14.3% | 1,418 | 9,937 |
| over 10,000 | 15.0% | 2,006 | 13,411 |
Embed this figure
26.3% of one-person companies publish an llms.txt, against 15.0% of companies above ten thousand staff. A one-person company is a Wix, GoDaddy or WordPress site and the platform wrote the file. A ten-thousand-person company has a hand-maintained web presence, and a new root file means a ticket, a review and an owner. The sites with the most content worth describing are the least likely to have described it.
548.4 thousand websites contradict themselves
3.1% of websites with an llms.txt, 548,441 of them, serve it beside a robots.txt that blocks every unlisted crawler. A further 2.3%, 406,952 websites, publish one and block GPTBot by name.
A file addressed to language models, served by a site whose robots.txt tells every crawler to go away. Nobody decided both. One of the two was written by a platform, and the owner never reconciled them. Until an AI company commits to reading llms.txt, it will keep being written by software and ignored by models, and the contradiction count will only grow.
Frequently asked questions
How many websites have an llms.txt?
11.4% of the web: 17,956,155 of the internet's 158,187,361 live websites. About two fifths of them were generated by Wix and GoDaddy, and 334,864 sites serve an llms.txt and no robots.txt at all.
Who writes most of the web's llms.txt?
Platforms. Wix generates one on 58.4% of the sites on its nameservers and GoDaddy on 24.4%. Shopify's storefront template is another 12.1% of websites with one, and three WordPress SEO plugins another 10.2%.
Is llms.txt used by AI companies?
No operator has committed to reading it for permissions, and it was never designed as a permissions file. It describes content. That is why the 723,109 websites carrying a Disallow-Training line achieve nothing.
What is Disallow-Training in llms.txt?
A directive that exists in no specification. 722,605 of the 723,109 websites that carry it share one identical first line, the template GoDaddy's site builder writes. No AI operator honours it.
How much of the adoption number is real?
64.2% of websites with an llms.txt serve one with headings and links over two hundred bytes. 19.9% serve one that is empty or belongs to a site that is not open. Most of the structured ones were still generated by a platform rather than written by the site owner.
Does publishing an llms.txt stop AI training?
No. It is a description, not a directive, and the format has no syntax for refusing anything. Part two covers what blocking an AI crawler actually looks like and how many sites do it.
Methodology and sources
Every figure comes from the StackScan September 2026 crawl, which covers the entire internet. We requested /llms.txt from 158,187,361 live websites in the September 2026 crawl and kept the 17,956,155 websites that returned one with a 200 status.
Classification is substring matching on the stored file text, which is sound for generator signatures and literal directive names: either the file contains Disallow-Training or it does not. That differs from parsing robots.txt rules, where group precedence matters, and the two methods are not mixed.
The DNS provider is read from each site's nameservers at crawl time, one provider per domain. Server country comes from the IP address that answered. Link counts come from Markdown link syntax across every llms.txt of two hundred bytes or more, and the destination chart excludes any link containing the site's own domain. Company size comes from StackScan's company records, matched to sites by website address. That match covers a minority of sites and leans towards businesses, so the company chart describes businesses rather than the whole web.
Two limits. This is one window rather than a trend, and a fixed panel is frozen so the next run can be compared with this one. And a website can fall into several rows of the contents chart at once, so those shares do not sum.
Every chart has a table view, and every figure can be cited with a link back to it. Figures are free to reuse with attribution. The format itself is described at llmstxt.org.