Discovery, not adoption: crawler retrieval of an llms.txt resource announced through robots.txt, across seven sites on one server, and what a server log can and cannot show

llms.txt is usually evaluated as an un-announced file placed at /llms.txt. I measured a different and narrower question: whether verified or network-attributed crawlers retrieve a Markdown resource when its URL is announced in robots.txt through a Sitemap: directive and in page-level HTML. Across six WordPress sites and 165 announced site-days, these crawlers requested /ai/llms.txt 1,549 times, 9.4 per site-day (95% interval 8.9 to 9.9). Ten un-referenced files in the same directory received zero requests across 60 file-months; one sibling referenced only by an HTML link tag received 13, all from OpenAI crawlers. The un-announced root /llms.txt, present on all seven sites, was requested twice. A seventh site with a root file only received 373 passing crawler requests and none for llms.txt. ClaudeBot accounted for 1,131 requests and retrieved the resource a median 8 seconds after its most recent same-site robots.txt request; it retrieved the site's XML sitemap, named in the same robots.txt, at the same rate. Because the referenced resource is Markdown rather than a conforming XML sitemap, this measures URL retrieval after robots.txt discovery, not support for llms.txt or valid sitemap parsing. These observations establish retrieval of an announced URL. They do not establish recognition of llms.txt, parsing, indexing, ranking, or use in AI-generated answers. Tell a crawler where a file is and it fetches the file. That is discovery, not adoption. Measurement note, version 1. The deposit archive contains the 8,336 crawler-claimed request rows with verdicts (client domains lettered, no human visitor addresses), the analysis scripts and their verbatim output, the figure, the timestamped preregistration for the two follow-up experiments described in Section 5, a data dictionary, and a SHA-256 manifest. Competing interests: the author owns Lefty Media Co and builds the plugin that announces the file measured here; see Section 8.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-17
DOI
https://doi.org/10.5281/zenodo.22814843
Primary Topic
Web Data Mining and Analysis
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Discovery, not adoption: crawler retrieval of an llms.txt resource announced through robots.txt, across seven sites on one server, and what a server log can and cannot show

Nathan Hall
Zenodo (CERN European Organization for Nuclear Research)
Web Data Mining and Analysis
preprint

Discovery, not adoption: crawler retrieval of an llms.txt resource announced through robots.txt, across seven sites on one server, and what a server log can and cannot show

Nathan Hall
preprint en

Abstract

llms.txt is usually evaluated as an un-announced file placed at /llms.txt. I measured a different and narrower question: whether verified or network-attributed crawlers retrieve a Markdown resource when its URL is announced in robots.txt through a Sitemap: directive and in page-level HTML. Across six WordPress sites and 165 announced site-days, these crawlers requested /ai/llms.txt 1,549 times, 9.4 per site-day (95% interval 8.9 to 9.9). Ten un-referenced files in the same directory received zero requests across 60 file-months; one sibling referenced only by an HTML link tag received 13, all from OpenAI crawlers. The un-announced root /llms.txt, present on all seven sites, was requested twice. A seventh site with a root file only received 373 passing crawler requests and none for llms.txt. ClaudeBot accounted for 1,131 requests and retrieved the resource a median 8 seconds after its most recent same-site robots.txt request; it retrieved the site's XML sitemap, named in the same robots.txt, at the same rate. Because the referenced resource is Markdown rather than a conforming XML sitemap, this measures URL retrieval after robots.txt discovery, not support for llms.txt or valid sitemap parsing. These observations establish retrieval of an announced URL. They do not establish recognition of llms.txt, parsing, indexing, ranking, or use in AI-generated answers. Tell a crawler where a file is and it fetches the file. That is discovery, not adoption. Measurement note, version 1. The deposit archive contains the 8,336 crawler-claimed request rows with verdicts (client domains lettered, no human visitor addresses), the analysis scripts and their verbatim output, the figure, the timestamped preregistration for the two follow-up experiments described in Section 5, a data dictionary, and a SHA-256 manifest. Competing interests: the author owns Lefty Media Co and builds the plugin that announces the file measured here; see Section 8.

Zenodo (CERN European Organization for Nuclear Research)
MediaTek (China) (CN)
Web Data Mining and Analysis
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Discovery, not adoption: crawler retrieval of an llms.txt resource announced through robots.txt, across seven sites on one server, and what a server log can and cannot show — Nathan Hall · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS