Discovery, not adoption: crawler retrieval of an llms.txt resource announced through robots.txt, across seven sites on one server, and what a server log can and cannot show

llms.txt is usually evaluated as an un-announced file placed at /llms.txt. I measured a different and narrower question: whether verified or network-attributed crawlers retrieve a Markdown resource when its URL is announced in robots.txt through a Sitemap: directive and in page-level HTML. Across six WordPress sites and 165 announced site-days, these crawlers requested /ai/llms.txt 1,549 times, 9.4 per site-day (95% interval 8.9 to 9.9). Ten un-referenced files in the same directory received zero requests across 60 file-months; one sibling referenced only by an HTML link tag received 13, all from OpenAI crawlers. The un-announced root /llms.txt, present on all seven sites, was requested twice. A seventh site with a root file only received 373 passing crawler requests and none for llms.txt. ClaudeBot accounted for 1,131 requests and retrieved the resource a median 8 seconds after its most recent same-site robots.txt request; it retrieved the site's XML sitemap, named in the same robots.txt, at the same rate. Because the referenced resource is Markdown rather than a conforming XML sitemap, this measures URL retrieval after robots.txt discovery, not support for llms.txt or valid sitemap parsing. These observations establish retrieval of an announced URL. They do not establish recognition of llms.txt, parsing, indexing, ranking, or use in AI-generated answers. Tell a crawler where a file is and it fetches the file. That is discovery, not adoption. Measurement note, version 1. The deposit archive contains the 8,336 crawler-claimed request rows with verdicts (client domains lettered, no human visitor addresses), the analysis scripts and their verbatim output, the figure, the timestamped preregistration for the two follow-up experiments described in Section 5, a data dictionary, and a SHA-256 manifest. Competing interests: the author owns Lefty Media Co and builds the plugin that announces the file measured here; see Section 8. Version 1.1 (2026-09-21): Section 4 now opens with the directory-link confound raised by John Mueller (Google) on 2026-09-17, that “everything llms.txt” directory sites sweep domains for the root file and link to it, so a crawler could arrive by following those links rather than the robots.txt Sitemap: line. The announced file measured here sits at /ai/llms.txt, a path those sweeps do not check; the root /llms.txt they would link received 2 passing crawler fetches against the announced file's 1,549. A Referer audit of every llms.txt request was added to the deposit (referer-check.mjs and its output): no fetch of either file by an identified crawler carried a third-party referer. It is reported as a supporting null, not proof, because these crawlers almost never send a Referer at all. NO COUNT IN THE PAPER CHANGES BETWEEN v1 AND v1.1. Both paper-v1-source.html and paper-v1.1-source.html are in the archive so that can be checked directly.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-21
DOI
https://doi.org/10.5281/zenodo.22880411
Primary Topic
Web Data Mining and Analysis
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Discovery, not adoption: crawler retrieval of an llms.txt resource announced through robots.txt, across seven sites on one server, and what a server log can and cannot show

Nathan Hall
Zenodo (CERN European Organization for Nuclear Research)
Web Data Mining and Analysis
preprint

Discovery, not adoption: crawler retrieval of an llms.txt resource announced through robots.txt, across seven sites on one server, and what a server log can and cannot show

Nathan Hall
preprint en

Abstract

llms.txt is usually evaluated as an un-announced file placed at /llms.txt. I measured a different and narrower question: whether verified or network-attributed crawlers retrieve a Markdown resource when its URL is announced in robots.txt through a Sitemap: directive and in page-level HTML. Across six WordPress sites and 165 announced site-days, these crawlers requested /ai/llms.txt 1,549 times, 9.4 per site-day (95% interval 8.9 to 9.9). Ten un-referenced files in the same directory received zero requests across 60 file-months; one sibling referenced only by an HTML link tag received 13, all from OpenAI crawlers. The un-announced root /llms.txt, present on all seven sites, was requested twice. A seventh site with a root file only received 373 passing crawler requests and none for llms.txt. ClaudeBot accounted for 1,131 requests and retrieved the resource a median 8 seconds after its most recent same-site robots.txt request; it retrieved the site's XML sitemap, named in the same robots.txt, at the same rate. Because the referenced resource is Markdown rather than a conforming XML sitemap, this measures URL retrieval after robots.txt discovery, not support for llms.txt or valid sitemap parsing. These observations establish retrieval of an announced URL. They do not establish recognition of llms.txt, parsing, indexing, ranking, or use in AI-generated answers. Tell a crawler where a file is and it fetches the file. That is discovery, not adoption. Measurement note, version 1. The deposit archive contains the 8,336 crawler-claimed request rows with verdicts (client domains lettered, no human visitor addresses), the analysis scripts and their verbatim output, the figure, the timestamped preregistration for the two follow-up experiments described in Section 5, a data dictionary, and a SHA-256 manifest. Competing interests: the author owns Lefty Media Co and builds the plugin that announces the file measured here; see Section 8. Version 1.1 (2026-09-21): Section 4 now opens with the directory-link confound raised by John Mueller (Google) on 2026-09-17, that “everything llms.txt” directory sites sweep domains for the root file and link to it, so a crawler could arrive by following those links rather than the robots.txt Sitemap: line. The announced file measured here sits at /ai/llms.txt, a path those sweeps do not check; the root /llms.txt they would link received 2 passing crawler fetches against the announced file's 1,549. A Referer audit of every llms.txt request was added to the deposit (referer-check.mjs and its output): no fetch of either file by an identified crawler carried a third-party referer. It is reported as a supporting null, not proof, because these crawlers almost never send a Referer at all. NO COUNT IN THE PAPER CHANGES BETWEEN v1 AND v1.1. Both paper-v1-source.html and paper-v1.1-source.html are in the archive so that can be checked directly.

Zenodo (CERN European Organization for Nuclear Research)
MediaTek (China) (CN)
Web Data Mining and Analysis
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.