BlogResearch

Does anything actually fetch your llms.txt? Here is what the server logs say

46 days of server logs from 39 business sites: named AI agents asked for /llms.txt 100 times, while five OpenAI and Anthropic fetchers made 195,245 requests to ordinary pages.

The Arrivl Team13 min read

Key takeaways

  • Named AI agents requested /llms.txt 100 times on the sites we measured between 1 August and 15 September 2026, spread over 46 days.
  • 21 of the sites we measured went those 46 days with zero requests for /llms.txt from any named agent.
  • 26 of the sites we measured serve an /llms.txt today that returns HTTP 200 with a text body, and 18 of them saw a request for it during the measurement window.
  • Between 25 August and 15 September 2026, on the sites carrying Cloudflare's verified-bot field, the 13 verified requests for /llms.txt came from PetalBot, Amazonbot, ChatGPT-User and bingbot, and GPTBot, OAI-SearchBot, ClaudeBot and Claude-User produced zero.
  • The five OpenAI and Anthropic fetchers in our data made 40 requests to /llms.txt against 195,245 requests to ordinary pages on the same sites, about one in 4,900.
  • Markdown twins of ordinary pages drew 8,179 requests on the sites we measured; one search engine on a single site accounted for 7,993 of them, and the other four sites with twin traffic sum to 186.
  • Ahrefs looked at 137,210 domains in its Web Analytics data for May 2026 and reported that 97% of the llms.txt files it found "received zero traffic in May".
  • Google states in its own AI features documentation, last updated July 2026, that "Google Search ignores them".
  • John Mueller of Google wrote on Bluesky on 17 June 2025 that "no AI system currently uses llms.txt."

What llms.txt is, and what the vendors say about reading it

Jeremy Howard of Answer.AI proposed the file on 3 September 2024: "We propose adding a /llms.txt file to websites that are designed for reading by language models, not just humans." He described it as "a markdown file that provides brief background information and guidance, along with links to markdown files (which can also link to external sites) providing more detailed information", on the reasoning that "Language models generally like to have information in a more concise form."

The second file most implementers ship, a root /llms-full.txt, comes from practice rather than from the proposal. The proposal text defines no root /llms-full.txt; its worked example names FastHTML's llms-ctx.txt and llms-ctx-full.txt. Cloudflare, Anthropic and Perplexity all serve an /llms-full.txt for their own documentation, which is how the convention spread.

Google is the one vendor that has answered the question directly, and it has answered it three times. Its AI features guide, last updated 10 July 2026, says: "You don't need to create new machine readable files, AI text files, markup, or Markdown to appear in Google Search (including its generative AI capabilities), as Google Search itself doesn't use them." The same page adds: "It's completely fine if you decide to create and maintain LLMS.txt files (or other similar files) for other services or systems that use these files. Doing so will neither harm nor help your site's visibility or rankings in Google Search, as Google Search ignores them." John Mueller of Google wrote on Reddit's r/TechSEO in April 2025, as carried by Search Engine Journal: "AFAIK none of the AI services have said they're using LLMs.TXT (and you can tell when you look at your server logs that they don't even check for it). To me, it's comparable to the keywords meta tag." Gary Illyes of Google said in July 2025 that Google does not support llms.txt and has no plans to, according to an attendee's report of the session rather than a recording.

The other vendors publish the file and say nothing about reading it. OpenAI's crawler documentation, read on 16 September 2026, names robots.txt as the control surface for GPTBot, OAI-SearchBot and ChatGPT-User, and we found no OpenAI statement about llms.txt; OpenAI's own docs site publishes one. Anthropic's crawler article, last updated 7 April 2026, says "Anthropic's Bots respect 'do not crawl' signals by honoring industry standard directives in robots.txt", and we found no Anthropic statement about llms.txt; Anthropic publishes an llms.txt and an llms-full.txt for its API docs. Perplexity publishes one for its docs and has said nothing about reading anyone else's. Cloudflare serves the pair for its developer docs and lists the files as an AI Index feature in its 26 September 2025 announcement.

The assistants tell their users the same thing. Asked in September 2026 what llms.txt is and whether a site needs one, Claude answered: "There's no published evidence it changes whether ChatGPT, Perplexity, or AI Overviews cite or recommend you", and added that no major AI system "has confirmed it fetches or uses llms.txt as a ranking or citation input." ChatGPT, given the same question, called it "an emerging community convention, not an official W3C/IETF standard."

That leaves one question open to measurement: what do the agents actually request?

How these numbers were measured

We counted every request from a named AI or search agent, classified by its user-agent string, on 39 business websites whose server logs we read, between 1 August and 15 September 2026, a span of 46 days. The named agents include GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-User, PerplexityBot, Googlebot, bingbot, Applebot, Amazonbot, Meta-ExternalAgent, PetalBot, CCBot and Bytespider. The sites are small and mid-size business sites in legal tech, online education, B2B software, e-commerce, services and media, and some of them have fewer than 46 days of data because install dates vary.

Cloudflare's verified-bot category exists on 9 of these sites and only from 25 August 2026. Numbers labelled verified use those 9 sites and that shorter window. Every other number on this page is what the user-agent string claimed. We compared four path groups: /llms.txt, /llms-full.txt, markdown twins (any path ending in .md, the markdown version of an ordinary page), and /agents.md. Query strings were removed before counting.

Who requests /llms.txt

agentrequestssites with ≥1 requestdays with ≥1 requestfirstlast
GPTBot229142026-08-042026-09-14
PetalBot162132026-08-072026-09-15
Meta-ExternalAgent121122026-08-032026-08-26
Googlebot11482026-08-012026-09-07
PerplexityBot8282026-08-032026-09-11
Claude-User7452026-08-032026-09-08
CCBot6652026-08-072026-09-14
OAI-SearchBot6342026-08-032026-08-28
Amazonbot4332026-09-082026-09-14
ChatGPT-User3232026-09-092026-09-12
bingbot3332026-08-142026-08-29
ClaudeBot2222026-08-102026-08-25
all named agents10018372026-08-012026-09-15

Sites we measured, 46 days, as claimed by user-agent.

The pattern is thin and wide. Twelve named agents asked for the file at least once, and the whole population of requests fits in 100 rows across 37 calendar days. The median site received zero requests for /llms.txt and the busiest received 20; twelve sites landed in the 1 to 5 range and six sites in the 6 to 20 range. GPTBot reached the most sites at 9, while Meta-ExternalAgent's 12 requests all landed on one site.

Two smaller behaviours sit in the same data. Where an agent came back to the same site's /llms.txt at all, which happened in 21 site-agent pairs over 59 intervals, the median gap between visits was 2.6 days. A third of all /llms.txt requests, 33%, fell inside the first seven days of a site's data, which reads as discovery at first crawl for at least part of the volume.

The neighbouring files are quieter still. /llms-full.txt drew 16 requests on five sites, led by PetalBot with 11, then CCBot with 2 and one each from Amazonbot, Applebot and Claude-User. /agents.md drew 21 requests on two sites, led by PetalBot with 10 and Googlebot with 5.

Cloudflare's verification changes the picture of who is really behind those user-agent strings.

pathagentverified requestssitesdays
/llms.txtPetalBot626
/llms.txtAmazonbot322
/llms.txtChatGPT-User323
/llms.txtbingbot111
/llms.txttotal134
markdown twinsClaudeBot2411
markdown twinsAmazonbot1511
markdown twinsOAI-SearchBot211
markdown twinstotal411

Cloudflare-verified requests, 9 sites, 25 August to 15 September 2026.

On those nine sites in those three weeks, GPTBot, OAI-SearchBot, ClaudeBot and Claude-User produced zero verified requests for /llms.txt. The few GPTBot requests we saw for the file could not be verified, and the same holds for the single PerplexityBot request there, which is four requests in total and far too few to generalize from. ChatGPT-User, the fetcher that acts when a person asks ChatGPT a question, produced 3 verified requests. Every other agent's /llms.txt requests on those sites carried verification when they happened.

Who requests the markdown twins instead

agentrequestssitesdays
bingbot7,993343
OAI-SearchBot73327
Amazonbot33312
ClaudeBot3033
Claude-User25114
GPTBot1513
PetalBot414
ChatGPT-User323
Meta-ExternalAgent312
all named agents8,179545

Sites we measured, 46 days, as claimed by user-agent.

The contrast is the finding. Named agents requested markdown twins 8,179 times and /llms.txt 100 times on the same sites over the same days. One site accounts for 7,993 of those twin requests, all of them from bingbot across 43 of the 46 days, so the honest comparison drops that site: the remaining four sites with twin traffic drew 186 twin requests against 100 requests for /llms.txt everywhere. Five sites had twin traffic at all, and most of the rest serve no twins to request.

The AI fetchers take twins in sweeps. ClaudeBot's 30 twin requests fell on three days, and 24 of them were a single same-day pass over 24 distinct pages on one site. Amazonbot behaved the same way, taking 15 pages on one site in one day. An agent that wants a site's content in markdown goes page by page through the pages, rather than to an index file.

Ordinary HTML pages dwarf both.

agentall requeststo /llms.txtto markdown twinsto ordinary pages
ChatGPT-User80,0933380,087
OAI-SearchBot47,08267347,003
ClaudeBot40,44723040,414
GPTBot18,271221518,233
Claude-User9,3527259,316

Sites we measured, 46 days, as claimed by user-agent.

For these five fetchers together, 40 requests went to /llms.txt and 195,245 went to ordinary pages, about one /llms.txt request for every 4,900 page requests. The file stayed under 0.12% of every one of these agents' traffic.

Who has one

We checked the files live on 16 September 2026. 26 of the sites we measured serve /llms.txt with HTTP 200 and a text body, and 20 serve /llms-full.txt. The file is usually present and usually unrequested: 26 sites serve it, and 18 saw a request for it in the window.

Published adoption numbers from others, each with its own sample, sit in a similar band. Rankability's tracker, dated 23 August 2026, reports that "8.7% of the top 1,000 websites publish an llms.txt file as of June 2026". ProGEO.ai's 31 March 2026 study of the Fortune 500 found that "only 7.4% of Fortune 500 companies - 37 in total - have implemented llms.txt". SE Ranking, reported by Search Engine Journal on 20 November 2025 across roughly 300,000 domains, "found llms.txt on 10.13% of domains" and measured no clear effect on AI citations. Originality.AI's tracker, updated 3 July 2026 over 3 million monitored sites, says "By May 2026, that number reached 36,120, an 8.8x increase in twelve months." In the Ahrefs study of 137,210 domains, about 38,000 carried a valid llms.txt, and Louise Linehan wrote on 15 June 2026 that "97% of those files received zero traffic in May. Nothing fetched them at all", with "OAI-SearchBot, PerplexityBot, and Claude's search crawler combined" making "only a couple of hundred fetches across thousands of sites".

Two single-site log studies land in the same place and should be read as single cases. Saaslinks reported on 3 August 2026, from one domain over 14 days, that crawlers "fetched /robots.txt 723 times and /llms.txt zero times". Weekerp reported on 16 August 2026 from two websites: "we recorded 68,759 AI bot requests", with "0 requests for llms.txt".

What this means for you

Should I publish an llms.txt?

Publish it if it costs you an hour, and budget nothing beyond that hour. Our measurement puts the file's demand at 100 requests across the sites we measured in 46 days, so the upside is small and so is the cost. The work that pays sits in the pages the agents already request tens of thousands of times.

I published one. Will ChatGPT read it?

The OpenAI fetchers asked for it rarely on the sites we measured. GPTBot made 22 requests to the file, OAI-SearchBot 6 and ChatGPT-User 3, while ChatGPT-User alone made 80,087 requests to ordinary pages on those sites. OpenAI's crawler documentation names robots.txt and says nothing about llms.txt.

What do the agents read instead?

They read your ordinary pages, and a few of them take the markdown twin of a page when one exists. On the sites we measured, five AI fetchers made 195,245 requests to ordinary pages. Claude, asked about serving markdown versions of pages in September 2026, put the priority the same way: "It's not a substitute for making individual pages fetchable in clean form."

Does llms.txt affect whether AI cites me?

No public measurement shows an effect, and the largest one looked for it. SE Ranking's crawl of roughly 300,000 domains, reported on 20 November 2025, found no clear effect on AI citations. Google states that having the file "will neither harm nor help your site's visibility or rankings in Google Search". Our own data counts requests, and a request is not a citation.

How do I check if anything fetched mine?

Read your own access log for the three paths and match on the user-agent. The counts land in minutes, and they answer the question for your site rather than for the average site.

How to check this on your own site

Start with the paths. Grep your access log for the three of them together, then split by user-agent:

grep -E ' /(llms\.txt|llms-full\.txt|agents\.md)' access.log
grep -E ' /[^ ]+\.md ' access.log

Strip query strings before you count, or one page will appear as six paths. Then apply two rules that keep the numbers honest. A user-agent string is a claim, so treat GPTBot in your log as something calling itself GPTBot until it is verified; behind Cloudflare, the verified-bot category tells you which requests the network itself confirmed. Read the denominator next to the count as well: the same log holds that agent's requests to your ordinary pages over the same days, and that is the number that puts an llms.txt count in scale.

FAQ

Is llms.txt an official standard? ChatGPT described it in September 2026 as "an emerging community convention, not an official W3C/IETF standard". Jeremy Howard published the proposal on 3 September 2024 and it remains a proposal.

Does Google use llms.txt? Google's AI features documentation, updated 10 July 2026, says "Google Search ignores them", and John Mueller wrote on Bluesky on 17 June 2025 that "no AI system currently uses llms.txt."

Which agents did ask for it? Twelve named agents asked at least once on the sites we measured, led by GPTBot with 22 requests on 9 sites and PetalBot with 16 requests on 2 sites.

Should I publish llms-full.txt too? It draws even less traffic: 16 requests on five of the sites we measured in 46 days, most of them from PetalBot. It also sits outside the original proposal, which defines no root /llms-full.txt.

Are markdown twins worth serving? Some agents take them in bulk when they exist. ClaudeBot fetched 24 twins on one site in a single day, and Amazonbot fetched 15 on one site in a single day, both verified by Cloudflare.

Arrivl reads these same requests out of your own logs, so you can see which agents asked for what on your site: see your own numbers.

llms.txtagentsmeasurement

See what AI agents do on your site — not in theory, on yours.

Start for free