Beyond similarity retrieval: governance fidelity, consistency, and minimality in long-term conversational memory
Abstract Conversational language model systems persist user constraints, preferences, identity, commitments, and assigned roles across turns. The dominant memory architecture stores this alongside ordinary episodic content in a vector index and retrieves the top- k items by similarity to the active query. We argue that this design is structurally inappropriate for the subset of items we call governing facts: facts whose binding force does not depend on their similarity to the current query. We define governance fidelity, the fraction of governing facts that reach the assembled context, and prove a small impossibility result: any retriever that ranks by a type-blind similarity score and returns a bounded top- k can be defeated by an adversarial topical-noise sequence. The remedy is type-conditional retrieval. We instantiate it as Stratified Context Reconstruction (SCR) in Satchel, a typed graph memory. Across 2880 sparse-and-hybrid and 2160 dense retrieval measurements, governance fidelity collapses from 1.0 to below 0.10 by N = 100 adversarial noise pairs for every similarity retriever tested; dense bi-encoders collapse as well, and a follow-up sweep over larger encoders (bge-large, e5-large) and a cross-encoder reranker confirm that the collapse is intrinsic to similarity ranking and is not avoided by stronger or larger encoders, which postpone but do not prevent it. SCR holds fidelity at 1.0 by construction (Cohen’s d = 1.06, p ≈ 2.6 + 10⁻ 25 ). We further show that type-conditional retrieval alone is insufficient once governing facts are revised over time: we prove a second impossibility—governance consistency cannot be guaranteed by any conflict-blind retriever, and the error grows linearly in revision depth—and give SCR-T, a temporally stratified, supersession-resolved retriever that attains both completeness and consistency by construction. A revision-depth benchmark confirms the prediction: naive SCR returns up to fifteen stale superseded facts per query and its governance minimality—the precision of the governing block—decays to 0.25, while SCR-T returns none and holds fidelity, consistency, and minimality jointly at 1.0. We frame these three as governance invariants against which conversational memory should be evaluated, alongside traceability as a structural property of typed memory. On four small open-weight LLMs, SCR improves constraint adherence over a TF-IDF RAG baseline by 17.59 percentage points. A further live-LLM experiment under the paper’s own adversarial noise, with retrieval and prompt formatting separated in a controlled four-arm design, shows that SCR’s downstream advantage is a high-noise phenomenon that grows with noise volume and is not an artefact of formatting.
Authors
- Rasheed Mohammad Nassr (ORCID: https://orcid.org/0000-0002-0800-428X)
- Haitham Mahmoud
- Oghenenefe Abeke
Institutions
- Birmingham City University (GB)
Publication Details
- Journal
- International Journal of Data Science and Analytics
- Published
- 2026-10-07
- DOI
- https://doi.org/10.1007/s41060-026-01290-8
- Primary Topic
- Topic Modeling
- Type
- article
- Field-Weighted Citation Impact
- 0.00