Spare capacity, narrow surface: the exposure record of a production agent-memory system
A production memory system, serving six agents for twelve weeks, never exposed 83.78% of its own corpus — and not for lack of room. The proactive channel delivered 583,763 slots against 67,187 chunks, capacity to show everything eight times, and showed 1,787 distinct chunks: 2.66%. This work measures the exposure surface of a memory system for agents: how many distinct items an agent in production actually receives, and which ones. It is a coordinate the field's benchmarks do not measure — they evaluate nDCG and recall over query sets, and none of those metrics sees what no query ever reached. The field's canonical survey (TMLR 2602.06052v4, 218 papers) does not record the term. What is measured Non-exposure is the result of policy, not a capacity limit. Under uniform serving the expected coverage would be 99.98%. The two channels of the surface freeze, for opposite reasons. The coverage channel exhausts 100% of an eligible pool of 108 chunks every day; the main one is determined by search traffic from months ago and does not respond to score adjustment. Exposure falls with collection size. The five types with more than a thousand chunks are squeezed into 10.7–27.0% exposure; of the ten small types, eight are above 32.5% and two are at zero. Size, however, is not separable from curation: no corpus variable measures the latter. What is not measured, and the paper says so There is no instrumented downstream outcome: nothing here establishes that non-exposure costs the agent value. The generalization of the mechanism is deductive — it holds for any ranker with lexicographic order and a bonus on the subordinate coordinate — and not empirical: it is one system. And of the two measured surfaces, only the brief is decided by the system; search is initiated by the agent and accounts for most of the recorded exposure. Instrument defects as part of the contribution Section 6 and Appendix E catalogue 17 instrument defects found during the work, eight of them altering numbers that were already written. They are in the paper because the rate at which a production measurement produces silent defects is itself a result — and because each correction names who found it and what would have prevented it. Among them: a value cited in five places without any artifact that contained it; three coefficients that lived only in stdout for four days; an artifact rewritten on every run being cited as a historical source; and a guard whose regular expression made it incapable of failing. The package Besides the manuscript, the deposit includes the verifier (claims_check.py, 19 guards that recompute the numerical claims against the artifacts), the six mechanical censuses, the 58 measurement artifacts, the 53 scripts that generate them and the three serving modules pinned by the sha256 of their blobs. The intent is that every number in the paper be recomputable by a third party without access to the system — and the package carries the artifacts, not only their hashes, because a corpus pinned by identifier was already lost in this same work, taking 280 adjudicated episodes with it. Relationship to the pre-registration. This is a completed observational study. The interventional trial it precedes is pre-registered at 10.5281/zenodo.22110203 and has not yet started: no randomized epoch exists, no arm has been assigned. The deviations between the public registration and what will apply to Epoch 1 are declared in DEVIATIONS-FOR-PAPER.md, including the most delicate one — the registration still declares open a designation defect that was resolved outside it. ⛔ Erratum 2026-09-07 — one script in the package is a record of an error, not an instrument The file measurement/dose2.mjs, inside scripts.zip, was refuted on 2026-08-27 and the number it produces is false. It is in the deposit as a record of the error, and nothing inside it says so — this erratum exists so that whoever opens the package knows before running it. What it produces: churn 0 at w ∈ {2.0 · 4.0 · 7.5}, with 19 boosts emitted. What the real pipeline produces: 11 / 15 / 17 states out of 350, monotonic, saturating in (4.0 ; 4.4]. Why: "production path" there was the right corpus with the ordering reimplemented — it goes through neither interleaveFresh nor pickDedup. A positive control over a reconstructed pipeline proves nothing about the pipeline. ⚠️ The aggravating factor: the file's own header asserts what was refuted — "now on the path that prod REALLY uses" (in the original Portuguese, "agora no caminho que prod REALMENTE usa") — and narrates a harness defect already fixed, which leads one to conclude that that version is sound. It is not. No number in this paper depends on it. The current texts (MANUSCRIPT.md §5.4 and DEVIATIONS-FOR-PAPER.md §1), both deposited, give 11 / 15 / 17. The risk is exclusively to whoever reruns the script taking it for an instrument. The caveat had existed since 08-27 in measurement/README.md, which is not part of the deposit — the artifact travelled and the caveat stayed behind. The files of a published record are immutable, so the correction comes through here. Detail: paper2-interventional/deposit/paperA/POST-PUBLISH.md in the repository. Erratum (2026-09-08) — the interventional trial has started. The section Relationship to the pre-registration above states that the trial pre-registered at 10.5281/zenodo.22110203 “has not yet started: no randomized epoch exists, no arm has been assigned”. That was true when this record was published and ceased to be on 2026-09-01. The 234 epochs were randomized prospectively in a single operation (ASSIGNMENT.json; drand quicknet beacon, round 31774052, seed SHA256(ascii(randomness_hex))), balanced 117 control / 39 w=2 / 39 w=4 / 39 w=7.5, and Epoch 1 entered active mode at 2026-09-01T10:25:39Z. No trial result is contained in this record, and no conclusion of this observational study depends on the state of the trial — the corrected statement is descriptive, not analytical. Two honesty notes on this correction. First: the defect is one of aging, not an error at the time of writing — a sentence about a live series, published as if it were an instant, became false through the passage of time. It was detected by an automatic verifier that compares the text's claims against the state of the artifacts (claims_check.py), not by rereading. Second: the files of this deposit remain immutable and were not altered; this erratum is appended to the description, which is metadata. The MANIFEST remains the proof of the published content and must not be recomputed. What changed in v1.1 (2026-10-05) The text above is the v1.0 description with its two dated errata, translated from the Portuguese published with v1.0; content and numbers are unchanged, and the Portuguese original stays on the v1.0 record. Version 1.1 is the manuscript translated to English and revised: the 2026-09-21 corrections, a 2026-10-04 audit of claims that had aged, and the reviews of 2026-10-05. Every change is recorded, with its source, in Appendix F of the manuscript; the 2026-10-05 round is Appendix F-5. Where the v1.0 description and v1.1 differ, v1.1 holds. Title. v1.0 was titled "Spare capacity, narrow surface: what a production agent-memory system actually surfaces". v1.1 is titled "Spare capacity, narrow surface: the exposure record of a production agent-memory system": the old ending used "actually" as an intensifier and repeated surface/surfaces. The title note at the top of the manuscript records the change. Recomputed with production code. Three quantities were recomputed rather than relabelled, each with a dated artifact. Access counterfactual (§4.3.2). The published 2/3/5 → 131/129/128 came from a script whose recency and importance default are not those of the production function. With calculateSalience the three chunks rank 1, 3 and 4 by salience among the 149 served chunks and fall to a tie at ranks 44–46 when the access component is zeroed (23–47 across instants and access-timing assumptions). Dedup was not replayed, so these ranks do not by themselves establish final served membership. "Beyond rank 100" and "exactly 52 places lower" are withdrawn. The claim that zeroing access takes the three out of the top-10 stands. out/SALIENCE-COUNTERFACTUAL-PROD-2026-10-05.json. Tie-break exposure (§5.7.1). Recounted on the comparator's key instead of the SQL pre-rank expression: 34–42 pairs at second resolution (0.59–0.73%) and 1,197 at day resolution. The earlier 68 / 269 / 917 / 1,656 remain cited as the pre-rank count. out/TIEBREAK-EXPOSURE-PROD-2026-10-05.json. Bonus against step (§4.4, §5.4). The 0.0946 and 1.79× used a severity multiplier of 0.5 that only the replay's summary applies; production applies 0.25. The bonus is 0.0473, 0.90× the largest adjacent gap, and the claim that the items crossed several positions is withdrawn. out/BONUS-VS-STEP-2026-10-05.json. Claims corrected or narrowed. Several sentences of the v1.0 description above are superseded: "never exposed 83.78% of its own corpus": the 56,288 live chunks (83.78%) have neither a brief-log record nor a positive search counter. That is a lower bound on non-delivery by the brief and by tracked search, not a measurement that they were never exposed; non-delivery across all agent-facing search is not established. The bound also assumes that the brief's log write, which is fail-open and records no failure, never failed: a failed write would deliver items with no row. "delivered 583,763 slots ... and showed 1,787 distinct chunks: 2.66%": the brief log records 583,763 selected slots and 1,635 distinct live chunks (2.43%); 1,787 counts the 152 that were served and later deleted. The log is written before the text renderer,
Authors
- Luiz Antonio Busnello (ORCID: https://orcid.org/0009-0007-5911-8141)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-05
- DOI
- https://doi.org/10.5281/zenodo.22181414
- Primary Topic
- Information Retrieval and Search Behavior
- Type
- preprint