Discovery, not adoption: crawler retrieval of an llms.txt resource announced through robots.txt, across seven sites on one server, and what a server log can and cannot show
llms.txt is usually evaluated as an un-announced file placed at /llms.txt. I measured a different and narrower question: whether verified or network-attributed crawlers retrieve a Markdown resource when its URL is announced in robots.txt through a Sitemap: directive and in page-level HTML. Across six WordPress sites and 165 announced site-days, these crawlers requested /ai/llms.txt 1,549 times, 9.4 per site-day (95% interval 8.9 to 9.9). Ten un-referenced files in the same directory received zero requests across 60 file-months; one sibling referenced only by an HTML link tag received 13, all from OpenAI crawlers. The un-announced root /llms.txt, present on all seven sites, was requested twice. A seventh site with a root file only received 373 passing crawler requests and none for llms.txt. ClaudeBot accounted for 1,131 requests and retrieved the resource a median 8 seconds after its most recent same-site robots.txt request; it retrieved the site's XML sitemap, named in the same robots.txt, at the same rate. Because the referenced resource is Markdown rather than a conforming XML sitemap, this measures URL retrieval after robots.txt discovery, not support for llms.txt or valid sitemap parsing. These observations establish retrieval of an announced URL. They do not establish recognition of llms.txt, parsing, indexing, ranking, or use in AI-generated answers. Tell a crawler where a file is and it fetches the file. That is discovery, not adoption. Measurement note, version 1. The deposit archive contains the 8,336 crawler-claimed request rows with verdicts (client domains lettered, no human visitor addresses), the analysis scripts and their verbatim output, the figure, the timestamped preregistration for the two follow-up experiments described in Section 5, a data dictionary, and a SHA-256 manifest. Competing interests: the author owns Lefty Media Co and builds the plugin that announces the file measured here; see Section 8.
Authors
- Nathan Hall
Institutions
- MediaTek (China) (CN)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-17
- DOI
- https://doi.org/10.5281/zenodo.22814844
- Primary Topic
- Web Data Mining and Analysis
- Type
- preprint