When Learning Hides Forgetting: How Endpoint Evaluation Masks Memory Corruption in Continuously Learning RAG Agents
Can a domain-specialized retrieval-augmented agent learn continuously from interaction memory without corrupting itself, and can standard evaluation practice see the answer? Quality-gated consolidation is the natural defence, and we test what it actually defends. We evaluate four memory conditions (no learning, verbatim accumulation, an embedding gate, and a hybrid embedding plus LLM-judge gate) across four dimensions (knowledge, strategy, personalization, self-correction) over a thirty-day interaction stream, under greedy decoding for bitwise reproducibility, with predictions and their falsification consequences registered before each run. We report two findings. First, standard practice hides this failure class three times over: verbatim accumulation raises the composite score while degrading factual knowledge; corruption is late-onset and non-monotone, so single-checkpoint evaluation misreads it; and corrupted answers co-occur with the correct memory retrieved, so presence-based probes certify health while the agent degrades. Second, gating does filter accumulation noise, but interference re-enters through the correction channel itself: gated, true-in-general corrections are applied outside their scope, and gated distillates that lose a correction's contrastive boundary collapse under retrieval perturbation. Verbatim storage, the baseline gating exists to beat, retains corrections more robustly than our gated distillation. A correction is a fact plus a boundary against an error, and consolidation that keeps the fact while dropping the boundary converts corrections back into interference. Continuous learning should therefore be evaluated per-dimension, per-checkpoint and end-to-end, and corrections should be consolidated with their boundaries and their provenance intact: quality gating governs what enters memory; it cannot govern how memory is applied.
Authors
- ChulHwee Joo
Institutions
- Valve (United States) (US)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-05
- DOI
- https://doi.org/10.5281/zenodo.22124210
- Primary Topic
- Topic Modeling
- Type
- preprint