Hashing Is Not Retrieval: A Three-Arm Study of Context Accumulation in Small-Model Code Cartography

Hashing Is Not Retrieval is a prospectively frozen three-arm evaluation of context accumulation in small-model multi-agent code cartography. It compares three ways for an orchestrator to hand work to a worker: 1. FULL HISTORY — retransmit prior worker outputs.2. OPAQUE REF — replace the previous output with an unretrievable SHA-256 digest.3. RETRIEVED — resolve the digest through a tenant-scoped content-addressed store and inject the recovered bytes. All arms receive identical source-derived evidence. The apparatus measures the exact message content sent to workers and records token counts, model and provider identifiers, generation IDs, usage, cost, retries, and retrieval proofs. A five-task gate preceded a fixed 50-task controlled benchmark, followed by a separate state-required evaluation across 24 deterministic fixtures and a five-scale Medusa case study. Key result: in the controlled benchmark, RETRIEVED reduced paired peak context by 38.3% relative to FULL HISTORY. The state-required follow-up shows that opaque references fail when state must be recovered, while verified retrieval preserves substantially more structural information. This is bounded-context evidence for automatic orchestrator-resolved state passing inside a specific two-worker code-cartography harness — not a universal claim about autonomous memory, long-context capability, or reasoning quality. The release includes the corrected author-identified preprint, executable apparatus, frozen results, task fixtures, telemetry, and verification instructions. Version 1.0.2 supplies the curated, verifier-passing reproducibility archive. It supersedes the uncurated archive from version 1.0.1 and includes the exact pinned apparatus, state-required fixtures, frozen results, and archive verifier. Submission status: This preprint was submitted to the NeurIPS 2026 Workshop on Long Context Foundation Models (LCFM). LCFM is a non-archival workshop; this Zenodo record is a public preprint and does not indicate acceptance or publication. Workshop information: https://longcontextfm.github.io/ Research website: https://abdulmoizahmed.com/researchConcept DOI (all versions): https://doi.org/10.5281/zenodo.22844754License: CC BY 4.0.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-19
DOI
https://doi.org/10.5281/zenodo.22844754
Primary Topic
Natural Language Processing Techniques
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Hashing Is Not Retrieval: A Three-Arm Study of Context Accumulation in Small-Model Code Cartography

Abdul Moiz Ahmed
Zenodo (CERN European Organization for Nuclear Research)
Natural Language Processing Techniques
article

Hashing Is Not Retrieval: A Three-Arm Study of Context Accumulation in Small-Model Code Cartography

Abdul Moiz Ahmed
article en

Abstract

Hashing Is Not Retrieval is a prospectively frozen three-arm evaluation of context accumulation in small-model multi-agent code cartography. It compares three ways for an orchestrator to hand work to a worker: 1. FULL HISTORY — retransmit prior worker outputs.2. OPAQUE REF — replace the previous output with an unretrievable SHA-256 digest.3. RETRIEVED — resolve the digest through a tenant-scoped content-addressed store and inject the recovered bytes. All arms receive identical source-derived evidence. The apparatus measures the exact message content sent to workers and records token counts, model and provider identifiers, generation IDs, usage, cost, retries, and retrieval proofs. A five-task gate preceded a fixed 50-task controlled benchmark, followed by a separate state-required evaluation across 24 deterministic fixtures and a five-scale Medusa case study. Key result: in the controlled benchmark, RETRIEVED reduced paired peak context by 38.3% relative to FULL HISTORY. The state-required follow-up shows that opaque references fail when state must be recovered, while verified retrieval preserves substantially more structural information. This is bounded-context evidence for automatic orchestrator-resolved state passing inside a specific two-worker code-cartography harness — not a universal claim about autonomous memory, long-context capability, or reasoning quality. The release includes the corrected author-identified preprint, executable apparatus, frozen results, task fixtures, telemetry, and verification instructions. Version 1.0.2 supplies the curated, verifier-passing reproducibility archive. It supersedes the uncurated archive from version 1.0.1 and includes the exact pinned apparatus, state-required fixtures, frozen results, and archive verifier. Submission status: This preprint was submitted to the NeurIPS 2026 Workshop on Long Context Foundation Models (LCFM). LCFM is a non-archival workshop; this Zenodo record is a public preprint and does not indicate acceptance or publication. Workshop information: https://longcontextfm.github.io/ Research website: https://abdulmoizahmed.com/researchConcept DOI (all versions): https://doi.org/10.5281/zenodo.22844754License: CC BY 4.0.

Zenodo (CERN European Organization for Nuclear Research)
National University of Computer and Emerging Sciences (PK)
Openalex Percentile: Top 8%
Natural Language Processing Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.