New Monthly plans from $39/mo. No contracts, cancel anytime. See pricing →

llms.txt Statistics 2026: What 17.9 Million Websites Publish

llms.txt has real adoption and the adoption figure is quoted everywhere. Almost nobody has looked at who wrote them or what is in them. These are the numbers, counted across the entire internet: how many sites have one, which platforms generated them, what they say, and how many contradict the site's own robots.txt.

Published 10 September 2026 · Updated 11 September 2026 · 13 min read · Part 3
Web crawler statistics 2026
  1. Part 1: Robots.txt statistics 2026
  2. Part 2: AI crawler statistics 2026
  3. Part 3: llms.txt statistics 2026
Melanie Cohen
Melanie writes about web infrastructure measurement at StackScan, working from the crawl that fetches robots.txt, ads.txt and llms.txt across the entire internet.
Key findings
  1. 11.4% of the web has an llms.txt: 17,956,155 of the internet's 158,187,361 live websites. That is far above every published estimate, and the rest of this list is why.
  2. Wix generated the llms.txt of 3.5 million websites and GoDaddy of 3.6 million, two fifths of the total between them. Another 2.2 million websites, 12.1%, serve Shopify's storefront template.
  3. 19.9% of websites with an llms.txt, 3,570,999 of them, serve one that is empty or describes a site that is not open: launching soon, store unavailable, or an HTML page served under the name.
  4. 723,109 websites use a Disallow-Training directive that appears in no specification. 722,605 of them are one identical line written by GoDaddy's site builder.
  5. Three WordPress SEO plugins wrote the llms.txt of 1.8 million websites, 10.2% of the total. 2.3 million websites serve one that uses the word "generated" about itself.
  6. The most linked destinations across every llms.txt are shop.app, dev.wix.com and shopify.com. The single most linked destination across the web's llms.txt is a platform's own property, not any site's content.
  7. 1.1 million websites serve an llms.txt over 200 bytes that contains no links at all. The median one has 5 links.
  8. 548.4 thousand websites publish an llms.txt and block every unlisted crawler in robots.txt. 407 thousand publish one and block GPTBot.
  9. 26.3% of one-person companies publish an llms.txt against 15.0% of companies above ten thousand staff. A standard pitched at publishers is being adopted by hosting defaults on the smallest sites.
17,956,155websites serve an llms.txt
11.4%of the web has one
19.9%serve one that is empty or belongs to a closed site
58.4%of sites on Wix nameservers have one

How many websites have an llms.txt

llms.txt was proposed in 2024 as a Markdown file at the root of a site that tells a language model what the site is and where its useful content lives. 11.4% of the web has one. We fetched /llms.txt from the entire internet, 158,187,361 live websites, and 17,956,155 returned one.

That is many times higher than the published estimates, which sample large publishers or count GitHub repositories. The gap is explained by who wrote them.

Who wrote them: Wix, GoDaddy and Shopify

Who wrote the web's llms.txt
Four generators account for most of the web's llms.txt. The site owner wrote the rest, or a placeholder did.
Who wrote the web's llms.txt Wix 3,452,053 generated for every published site GoDaddy 3,642,299 site builder, with a directive that does not exist Shopify 2,169,403 storefront template addressed to shopping agents WordPress SEO plugins 1,839,007 All in One SEO, Yoast, Rank Math 17,956,155 websites serve an llms.txt, the entire internet About 62 in 100 came from one of these four generators. A file can carry two signatures, so the share is approximate.
Embed this figure
<a href="https://www.stackscan.com/blog/llms-txt-statistics#fig-who-wrote"><img src="https://content.stackscan.com/charts/3-who-wrote.webp" alt="Who wrote the web's llms.txt" width="880" style="max-width:100%"></a> <p>Source: <a href="https://www.stackscan.com/blog/llms-txt-statistics#fig-who-wrote">StackScan llms.txt analysis</a></p>
llms.txt adoption by DNS provider
Share of websites on each provider's nameservers that serve an llms.txt, and the number of websites. Providers with fewer than 300,000 matched sites are left out.
llms.txt adoption by DNS providerWix58.4% (3,452,053)GoDaddy24.4% (3,642,299)Google14.5% (835,005)AWS Route 539.4% (135,568)Namecheap7.4% (287,318)Cloudflare7.1% (1,391,646)OVH5.9% (80,600)WordPress3.4% (39,178)NS13.0% (20,940)Squarespace1.8% (34,566)
Embed this figure
<a href="https://www.stackscan.com/blog/llms-txt-statistics#fig-adoption"><img src="https://content.stackscan.com/charts/3-adoption.webp" alt="llms.txt adoption by DNS provider" width="880" style="max-width:100%"></a> <p>Source: <a href="https://www.stackscan.com/blog/llms-txt-statistics#fig-adoption">StackScan llms.txt analysis</a></p>

Sites on Wix nameservers serve an llms.txt at 58.4%. Wix generates the file for its site builder customers, and among Wix sites that also publish a robots.txt, the closer proxy for a published site, the rate is 90.7%.

GoDaddy's site builder generates one on 24.4% of the sites on GoDaddy nameservers. Squarespace, at 1.8%, and WordPress.com, at 3.4%, made the opposite decision, and their customers have almost none.

Who wrote the web's llms.txt
Websites on the provider's nameservers, or websites whose llms.txt carries the signature. A website can appear in more than one row.
Who wrote the web's llms.txtMentions Wix4,328,723GoDaddy (DNS)3,642,299Wix (DNS)3,452,053Says "generated"2,266,801Mentions Shopify2,169,403Cloudflare (DNS)1,391,646All in One SEO1,118,229Google (DNS)835,005Yoast SEO523,198Namecheap (DNS)287,318Rank Math SEO197,580
Embed this figure
<a href="https://www.stackscan.com/blog/llms-txt-statistics#fig-platforms"><img src="https://content.stackscan.com/charts/3-platforms.webp" alt="Who wrote the web's llms.txt" width="880" style="max-width:100%"></a> <p>Source: <a href="https://www.stackscan.com/blog/llms-txt-statistics#fig-platforms">StackScan llms.txt analysis</a></p>

Wix wrote the llms.txt of 3,452,053 websites and GoDaddy of 3,642,299, 19.2% and 20.3% of the total. 2.2 million websites, 12.1%, serve Shopify's generated version, which addresses AI shopping agents rather than search engines.

The three WordPress SEO plugins, All in One SEO, Yoast and Rank Math, wrote 1.8 million between them, 10.2%. 2.3 million websites serve one that describes itself with the word "generated".

This is the WordPress finding from part one arriving in a younger file. The most widely deployed crawl instruction on the internet is a default nobody chose, and the most widely deployed AI instruction now is too. Any llms.txt adoption figure measures which platforms shipped a generator, not what publishers want, and the number will jump every time a large platform switches one on.

What websites put in it

What websites put in llms.txt
Websites and share of the 17.9 million that serve one. A website can be several of these at once.
What websites put in llms.txtEmpty1,723,099Under 200 bytes3,591,513Headings and links, 200 bytes or more11,521,675Says the site is launching soon1,224,112Says the store is unavailable429,921HTML or a script, not text193,867
Classification is substring matching on the stored text.
Embed this figure
<a href="https://www.stackscan.com/blog/llms-txt-statistics#fig-quality"><img src="https://content.stackscan.com/charts/3-quality.webp" alt="What websites put in llms.txt" width="880" style="max-width:100%"></a> <p>Source: <a href="https://www.stackscan.com/blog/llms-txt-statistics#fig-quality">StackScan llms.txt analysis</a></p>
A drawer of typed index cards from a card catalogue
An llms.txt is a card catalogue for a website: a description of what is inside, not a lock on the door. Most of the internet's were typed by a platform. Photo: Michael Holley, public domain, via Wikimedia Commons.

19.9% of websites with an llms.txt, 3,570,999 of them, serve one that is empty or belongs to a site that is not open: nothing at all, a parked domain announcing a launch, a closed storefront, or an HTML page served under the wrong name.

The 429.9 thousand "Store Unavailable" versions are Shopify's placeholder for a password-protected or suspended shop, faithfully describing a store nobody can enter. GoDaddy serves an llms.txt of zero bytes on 198.5 thousand websites.

The counterweight matters. 11,521,675 websites, 64.2%, serve one that is genuinely structured: Markdown headings, links, two hundred bytes or more. A clear majority is a real attempt at the format, even where a platform made the attempt. The adoption figure everyone quotes counts the placeholders and the templates too.

723,109 websites use a directive that does not exist

723,109websites use a Disallow-Training line
722,605of those are one identical template line
791,861websites carry robots.txt syntax
548,441also block every unlisted crawler in robots.txt

4.4% of websites with an llms.txt, 791,861 of them, have a User-agent: line in it. That is robots.txt syntax in a file specified as Markdown, and it does nothing. 723,109 of them go further and carry a Disallow-Training: directive. It appears in no specification and is honoured by no AI operator.

Those owners believe they have opted out of model training. They have not, in the same way the 54,109 sites in part two blocking Anthropic's retired user agents have not. This group is thirteen times larger.

It is also not a snippet that spread from blog post to blog post. 722,605 of those 723,109 websites begin the file with the identical line, User-agent: * Allow: / Disallow-Training: / Sitemap: /sitemap.xml.

By nameservers, 369.8 thousand of those sites are on GoDaddy, and most of the rest are on the two networks that host GoDaddy's website builder. The directive that does not exist was written by GoDaddy's builder once and shipped to every customer who published a site. The demand behind it is real: website owners want a permissions file for AI and reached for the nearest thing. Nobody on the AI side has given them one.

Directives that do nothing, and other things websites put in llms.txt
What the file containsFilesShare of llms.txt
A User-agent line791,8614.4%
Disallow-Training723,1094.0%
of which one identical template line722,6054.0%
Commercial-use: allowed28,6120.2%
Lorem ipsum203,3601.1%
The word "please"2,541,21414.2%
Chinese characters623,9463.5%
Cyrillic186,9141.0%
A link to llms-full.txt138,9110.8%
A link to llmstxt.org12,2430.1%
Over 1 MB4,4260.0%
Substring matching on the stored text.
Embed this figure
<a href="https://www.stackscan.com/blog/llms-txt-statistics#fig-wrong"><img src="https://content.stackscan.com/charts/3-wrong.webp" alt="Directives that do nothing, and other things websites put in llms.txt" width="880" style="max-width:100%"></a> <p>Source: <a href="https://www.stackscan.com/blog/llms-txt-statistics#fig-wrong">StackScan llms.txt analysis</a></p>

A second invented set, Commercial-use: allowed and Research-use: allowed, appears on 28.6 thousand websites, another generator's template. 203.4 thousand websites serve lorem ipsum in it, the placeholder text of a template nobody finished, handed to every language model that asks. Only 12.2 thousand link to the specification they follow.

A link is the whole point of the format. Across every llms.txt of two hundred bytes or more we counted 161,904,107 of them, a mean of 12.8 per website and a median of 5. Most are a very short map, and 1,100,303 websites serve one over two hundred bytes with no links whatsoever.

The most linked destinations in the web's llms.txt, excluding each site's own domain
Counted across every Markdown link in every llms.txt of 200 bytes or more
The most linked destinations in the web's llms.txt, excluding each site's own domainshop.app6,365,583dev.wix.com4,312,752ucp.dev2,122,452www.shopify.com2,121,703www.facebook.com405,857www.instagram.com334,102beacons.ai271,480docs.github.com246,302github.com179,417www.rediff.com155,790support.myclickfunnels.com140,760www.linkedin.com124,691
Counted across every llms.txt of 200 bytes or more. Any link containing the site's own domain is excluded.
Embed this figure
<a href="https://www.stackscan.com/blog/llms-txt-statistics#fig-targets"><img src="https://content.stackscan.com/charts/3-targets.webp" alt="The most linked destinations in the web's llms.txt, excluding each site's own domain" width="880" style="max-width:100%"></a> <p>Source: <a href="https://www.stackscan.com/blog/llms-txt-statistics#fig-targets">StackScan llms.txt analysis</a></p>

Strip out every link to the site's own domain and rank what is left. The top of the chart is not a publisher's content. Shopify's shop.app and shopify.com and Wix's developer documentation at dev.wix.com account for tens of millions of links inside llms.txt that are supposed to describe somebody else's website.

The single most linked destination across 18 million websites is a platform's own property. A file meant to guide language models to a site's best pages is, at scale, guiding them to the platform's checkout.

Adoption by company size

llms.txt adoption by company size
Share of sites in each headcount band that serve one, and the number of sites
llms.txt adoption by company size1 person26.3% (216,123)2 to 1025.5% (1,077,306)11 to 5022.6% (505,146)51 to 20020.2% (143,326)201 to 50018.1% (37,151)501 to 1,00017.2% (12,753)1,001 to 5,00015.3% (8,563)5,001 to 10,00014.3% (1,418)over 10,00015.0% (2,006)
Only sites matched to a company appear here.
Embed this figure
<a href="https://www.stackscan.com/blog/llms-txt-statistics#fig-headcount"><img src="https://content.stackscan.com/charts/3-headcount.webp" alt="llms.txt adoption by company size" width="880" style="max-width:100%"></a> <p>Source: <a href="https://www.stackscan.com/blog/llms-txt-statistics#fig-headcount">StackScan llms.txt analysis</a></p>

26.3% of one-person companies publish an llms.txt, against 15.0% of companies above ten thousand staff. A one-person company is a Wix, GoDaddy or WordPress site and the platform wrote the file. A ten-thousand-person company has a hand-maintained web presence, and a new root file means a ticket, a review and an owner. The sites with the most content worth describing are the least likely to have described it.

548.4 thousand websites contradict themselves

3.1% of websites with an llms.txt, 548,441 of them, serve it beside a robots.txt that blocks every unlisted crawler. A further 2.3%, 406,952 websites, publish one and block GPTBot by name.

A file addressed to language models, served by a site whose robots.txt tells every crawler to go away. Nobody decided both. One of the two was written by a platform, and the owner never reconciled them. Until an AI company commits to reading llms.txt, it will keep being written by software and ignored by models, and the contradiction count will only grow.

Frequently asked questions

How many websites have an llms.txt?

11.4% of the web: 17,956,155 of the internet's 158,187,361 live websites. About two fifths of them were generated by Wix and GoDaddy, and 334,864 sites serve an llms.txt and no robots.txt at all.

Who writes most of the web's llms.txt?

Platforms. Wix generates one on 58.4% of the sites on its nameservers and GoDaddy on 24.4%. Shopify's storefront template is another 12.1% of websites with one, and three WordPress SEO plugins another 10.2%.

Is llms.txt used by AI companies?

No operator has committed to reading it for permissions, and it was never designed as a permissions file. It describes content. That is why the 723,109 websites carrying a Disallow-Training line achieve nothing.

What is Disallow-Training in llms.txt?

A directive that exists in no specification. 722,605 of the 723,109 websites that carry it share one identical first line, the template GoDaddy's site builder writes. No AI operator honours it.

How much of the adoption number is real?

64.2% of websites with an llms.txt serve one with headings and links over two hundred bytes. 19.9% serve one that is empty or belongs to a site that is not open. Most of the structured ones were still generated by a platform rather than written by the site owner.

Does publishing an llms.txt stop AI training?

No. It is a description, not a directive, and the format has no syntax for refusing anything. Part two covers what blocking an AI crawler actually looks like and how many sites do it.

Methodology and sources

Every figure comes from the StackScan September 2026 crawl, which covers the entire internet. We requested /llms.txt from 158,187,361 live websites in the September 2026 crawl and kept the 17,956,155 websites that returned one with a 200 status.

Classification is substring matching on the stored file text, which is sound for generator signatures and literal directive names: either the file contains Disallow-Training or it does not. That differs from parsing robots.txt rules, where group precedence matters, and the two methods are not mixed.

The DNS provider is read from each site's nameservers at crawl time, one provider per domain. Server country comes from the IP address that answered. Link counts come from Markdown link syntax across every llms.txt of two hundred bytes or more, and the destination chart excludes any link containing the site's own domain. Company size comes from StackScan's company records, matched to sites by website address. That match covers a minority of sites and leans towards businesses, so the company chart describes businesses rather than the whole web.

Two limits. This is one window rather than a trend, and a fixed panel is frozen so the next run can be compared with this one. And a website can fall into several rows of the contents chart at once, so those shares do not sum.

Every chart has a table view, and every figure can be cited with a link back to it. Figures are free to reuse with attribution. The format itself is described at llmstxt.org.