AI search news ·
Do AI crawlers actually read llms.txt? DEJAN logged 4 big-three fetches against 1,150 for robots.txt in 30 days — our own ledger says 0 against 27
On the evidence of two independent server logs, almost never. DEJAN's Dan Petrovic published a census of every fetch of dejan.ai's three machine-readable files between 27 July and 15 August 2026: counting only Google, Anthropic and OpenAI crawlers, robots.txt was fetched 1,150 times, the site's OKF content bundle 23 times, and llms.txt 4 times — three by Googlebot, one by ClaudeBot, zero by any OpenAI agent. We opened our own crawler-hit ledger, live since 2026-08-17 01:48Z, and read the same three files over its first 39 hours: robots.txt 65 fetches by bot-shaped user agents, 27 of them from the same three vendors; llms.txt one fetch, from an agent that matches no recognised AI or search crawler; the sitemap 12. Google's John Mueller said it flatly in June 2025 — "FWIW no AI system currently uses llms.txt" — and fourteen months later two request logs, one large and one small, agree with him.
DEJAN's census, as published
The post (2026-08-15, Dan Petrovic; DEJAN is an AI-SEO consultancy, so this is a practitioner's own log, not a neutral lab — read it as such) counts fetches from dejan.ai's request log of three files: `/robots.txt`, `/llms.txt` (the llmstxt.org index for language models) and the site's OKF bundle at `/okf/`, its full machine-readable text. Method as stated: agents grouped by published User-Agent token; "Browser" is a normal browser UA, "Unrecognised" matches no known crawler or browser; the window is the 30 days ending 15 August, but logging began 27 July, so it holds 20 days of data; and the counts are fetches, "not renders, citations, or answer appearances." It was carried by Search Engine Roundtable's daily recap on August 17.
| File (dejan.ai, 30 d to 2026-08-15) | Anthropic | OpenAI | Big three | All agents | |
|---|---|---|---|---|---|
| robots.txt | 243 | 608 | 299 | 1,150 | 13,442 |
| OKF bundle (/okf/) | 11 | 1 | 11 | 23 | 2,841 |
| llms.txt | 3 | 1 | 0 | 4 | 157 |
Google is Googlebot plus GoogleOther and Google-CloudVertexBot; Anthropic is ClaudeBot (445 of the 608 robots.txt fetches), Claude-User (163) and anthropic-ai; OpenAI is OAI-SearchBot (295), GPTBot (1) and ChatGPT-User (3). The 157 llms.txt fetches from all agents came from 78 IP addresses and 13 agents: Browser 85 (54.1%), Unrecognised 40 (25.5%), curl 13 (8.3%), AhrefsBot 4, HeadlessChrome 4, Googlebot 3, Meta-ExternalAgent 2, and one each from ClaudeBot, MJ12bot, SemrushBot, Applebot, Bingbot and Amazonbot. Two details in the post are the ones to keep: the three Googlebot fetches of llms.txt "were followed by no page request from the same IP addresses," and OpenAI's agents — the ones that fetched robots.txt 299 times and the content bundle 11 times — never requested llms.txt once.
Our own ledger over its first 39 hours
Our crawler-hit ledger records every request whose user agent looks bot-shaped (matches `bot|crawl|spider|slurp|fetch|-user|…`) at the edge and classifies it against a self-declared-token table; browsers and curl are not recorded, so it is a crawler-only view and its window is short — it was born 2026-08-17 01:48Z and we read it at 16:44Z on the 18th. Our `/llms.txt` has been live since 2026-06-27 and is linked from nowhere on the site, in robots.txt or in the sitemap; a crawler finds it only by convention.
| File (askedabout.com, 39 h to 2026-08-18 16:44Z) | Anthropic | OpenAI | Big three | All bot-shaped agents | |
|---|---|---|---|---|---|
| robots.txt | 16 (Googlebot) | 1 (Claude-User) | 10 (OAI-SearchBot) | 27 | 65 (+ PerplexityBot 1, bingbot 3, YandexBot 5, 29 unrecognised) |
| sitemap.xml | 4 (Googlebot) | 0 | 1 (GPTBot) | 5 | 12 (+ bingbot 3, 4 unrecognised) |
| llms.txt | 0 | 0 | 0 | 0 | 1 (unrecognised bot-shaped agent, 2026-08-18 11:57Z) |
Same shape at a hundredth of the scale: robots.txt is read continuously (OAI-SearchBot alone came for it ten times in 39 hours, Googlebot sixteen), the sitemap is read, and llms.txt is not — one fetch, from nothing we recognise. We will keep the count running; if any of ChatGPT-User, OAI-SearchBot, GPTBot, ClaudeBot, Claude-User or PerplexityBot fetches `/llms.txt` in the next 30 days, we will say so with the timestamp. The linked-from-nowhere caveat cuts both ways: it means our zero is partly a discovery zero — but robots.txt and the sitemap are also found by convention, and they were found.
What this means for businesses that care about AI visibility
- llms.txt is a $0 hour of work with no measured return; robots.txt is the file every AI agent actually reads. Both logs show the AI vendors' agents fetching robots.txt hundreds of times and llms.txt approximately never. If you have one hour, spend it confirming robots.txt allows the agents you want — OAI-SearchBot, ChatGPT-User, PerplexityBot, ClaudeBot — and remember that user-triggered fetchers may not honour it either way.
- The people fetching your llms.txt are people. More than half of DEJAN's llms.txt requests carried a browser user agent and another 8% were curl — SEOs and tool vendors checking whether the file exists. That is who a listing in llms.txt reaches today, and it is not who decides whether ChatGPT names your business.
- What the engines do read is the page itself, in the plainest form you serve it. RESONEO's retrieval study found ChatGPT's own index snippets a page from ~202 characters starting at its H1 when it reads a page at all; our four-run corpus finds 92% of cited URLs are deep pages and 92% of cited slugs carry a question word. A machine index of your site is not the lever; a page that answers the buyer's question is. Whether the engines then name you when a buyer asks is the number our free 60-second check reads.
Do ChatGPT, Claude or Google actually fetch llms.txt?
Rarely to never, on the two server logs we have. In DEJAN's 30-day census to 2026-08-15, Google, Anthropic and OpenAI crawlers fetched dejan.ai's llms.txt 4 times (Googlebot 3, ClaudeBot 1, OpenAI 0) against 1,150 fetches of robots.txt; on askedabout.com's ledger over its first 39 hours, those vendors fetched robots.txt 27 times and llms.txt zero times. Google's John Mueller wrote in June 2025 that no AI system currently uses llms.txt.
Who is fetching llms.txt files, then?
Mostly humans and tools. Of 157 llms.txt fetches from all agents on dejan.ai, 85 (54.1%) carried a browser user agent, 40 were unrecognised agents and 13 were curl; SEO crawlers (AhrefsBot, SemrushBot, MJ12bot) and one-off fetches from Applebot, Bingbot, Amazonbot and Meta-ExternalAgent made up most of the rest.
Should I still add an llms.txt file?
It costs nothing and does no harm, but neither log shows an AI vendor's agent using it, and the three Googlebot fetches DEJAN saw were followed by no page requests. Spend the hour on robots.txt allow rules for the AI agents you want and on pages that answer buyer questions in their first 200 characters — those are the things the logs and the citation data show being read.
See your number
See which businesses AI names when your client's buyers ask.
Running this for clients? The $249 agency 5-pack audits five businesses, white-labeled.
Who runs this
- Built and operated by Sensara LLC, Atlanta, Georgia — about us and how the audit works.
- See what the report looks like before you run anything — score per engine, the competitors AI names instead of you, and a fix plan.
- We run the same audit on ourselves every week and publish the result: in the latest run AI named AskedAbout in 1 of 144 answers. We report our own numbers the way we report yours.