Retrieval, Not Generation, on a Microcontroller: Stratified Evaluation of Offline Doctrinal Search in 16 MiB

Language models now run on cheap microcontrollers, but the systems demonstrated there generate text and cannot answer a question or name a source. We ask the complementary question: what does retrieval cost, and deliver, on the same class of hardware. We build a hybrid lexical and semantic retriever over 1,337 passages drawn from four field manuals, with the embedding table resident in flash and the entire query path executing on an ESP32-S3 with 512 KB of SRAM, and we evaluate it with an instrument built for the case that matters operationally: a stratified set of 95 question groups, each phrased at four increasing levels of lexical distance from the manual, written and annotated by a military cadet rather than generated by a language model. The device returns the correct chapter among its ten results in 86.9% of queries (95% CI [81.9, 91.2]), in 11.45 milliseconds, using 9.83 MiB of the 16 MiB of flash, and it is verified bit-exact against a host reference on 50 of 50 fixtures. Exact-passage accuracy, however, collapses with lexical distance: hit@3 falls from 57.5% [46.2, 68.8] when the query is phrased naturally to 23.8% [15.0, 33.8] when it is phrased functionally, intervals that do not overlap. The failure is not blindness but resolution. Even at the scenario level the correct passage sits at median rank 9 of 1,337, inside a candidate cluster whose members average pairwise cosine 0.505 against 0.256 for two random passages, so what survives retrieval is twenty paragraphs about the same thing, and separating them requires procedural knowledge that no model of this size carries. Table size saturates below the device’s memory budget, and a transformer encoder that does not fit the hardware exceeds the deployed static table by 2.5 points overall, significantly only on scenario queries. We report five negative results supporting that diagnosis, each with its mechanism identified, under a protocol with thresholds fixed in advance and a test split opened once.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-06
DOI
https://doi.org/10.5281/zenodo.23188690
Primary Topic
Information Retrieval and Search Behavior
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

Retrieval, Not Generation, on a Microcontroller: Stratified Evaluation of Offline Doctrinal Search in 16 MiB

Vladin-Antonio Foca
Zenodo (CERN European Organization for Nuclear Research)
Information Retrieval and Search Behavior
preprint

Retrieval, Not Generation, on a Microcontroller: Stratified Evaluation of Offline Doctrinal Search in 16 MiB

Vladin-Antonio Foca
preprint en

Abstract

Language models now run on cheap microcontrollers, but the systems demonstrated there generate text and cannot answer a question or name a source. We ask the complementary question: what does retrieval cost, and deliver, on the same class of hardware. We build a hybrid lexical and semantic retriever over 1,337 passages drawn from four field manuals, with the embedding table resident in flash and the entire query path executing on an ESP32-S3 with 512 KB of SRAM, and we evaluate it with an instrument built for the case that matters operationally: a stratified set of 95 question groups, each phrased at four increasing levels of lexical distance from the manual, written and annotated by a military cadet rather than generated by a language model. The device returns the correct chapter among its ten results in 86.9% of queries (95% CI [81.9, 91.2]), in 11.45 milliseconds, using 9.83 MiB of the 16 MiB of flash, and it is verified bit-exact against a host reference on 50 of 50 fixtures. Exact-passage accuracy, however, collapses with lexical distance: hit@3 falls from 57.5% [46.2, 68.8] when the query is phrased naturally to 23.8% [15.0, 33.8] when it is phrased functionally, intervals that do not overlap. The failure is not blindness but resolution. Even at the scenario level the correct passage sits at median rank 9 of 1,337, inside a candidate cluster whose members average pairwise cosine 0.505 against 0.256 for two random passages, so what survives retrieval is twenty paragraphs about the same thing, and separating them requires procedural knowledge that no model of this size carries. Table size saturates below the device’s memory budget, and a transformer encoder that does not fit the hardware exceeds the deployed static table by 2.5 points overall, significantly only on scenario queries. We report five negative results supporting that diagnosis, each with its mechanism identified, under a protocol with thresholds fixed in advance and a test split opened once.

Zenodo (CERN European Organization for Nuclear Research)
Nicolae Bălcescu Land Forces Academy (RO)
Information Retrieval and Search Behavior
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.