A Screened Benchmark Dataset and Validity Study for Organizational Memory in Agent Harnesses
Organizations now use agents built on different models and software environments. Shared knowledge therefore cannot live in vendor-owned model weights; it must live in a memory layer accessible to every agent. We reviewed ten published memory benchmarks and found none tests whether memory preserves knowledge with the right scope, authority, and freshness. We introduce memory-bench, a benchmark and validity study for organizational memory. It simulates three software organizations with event streams, hidden belief-state ledgers, and behavioral work tasks. Each probe has a counterfactual twin in which the tested fact changes. An instance counts only if both versions pass. Across three released seeds, 371 instances passed screening, although all fell below the prespecified threshold of 135. A memoryless worker earned pair credit on 4 of 125 instances. Probe validity varied by worker, two harnesses with identical weights agreed on only 65% of outcomes, and judge-human agreement reached κ = 0.537, below the required 0.75. We retain these failures, document the loss of the pinned worker, and provide a re-anchoring procedure. The paper reports no comparison between memory systems.
Authors
- Mehul Srivastava (ORCID: https://orcid.org/0009-0008-1031-304X)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-19
- DOI
- https://doi.org/10.5281/zenodo.22838320
- Primary Topic
- Team Dynamics and Performance
- Type
- preprint