A Screened Benchmark Dataset and Validity Study for Organizational Memory in Agent Harnesses

Organizations now use agents built on different models and software environments. Shared knowledge therefore cannot live in vendor-owned model weights; it must live in a memory layer accessible to every agent. We reviewed ten published memory benchmarks and found none tests whether memory preserves knowledge with the right scope, authority, and freshness. We introduce memory-bench, a benchmark and validity study for organizational memory. It simulates three software organizations with event streams, hidden belief-state ledgers, and behavioral work tasks. Each probe has a counterfactual twin in which the tested fact changes. An instance counts only if both versions pass. Across three released seeds, 371 instances passed screening, although all fell below the prespecified threshold of 135. A memoryless worker earned pair credit on 4 of 125 instances. Probe validity varied by worker, two harnesses with identical weights agreed on only 65% of outcomes, and judge-human agreement reached κ = 0.537, below the required 0.75. We retain these failures, document the loss of the pinned worker, and provide a re-anchoring procedure. The paper reports no comparison between memory systems.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-19
DOI
https://doi.org/10.5281/zenodo.22838321
Primary Topic
Team Dynamics and Performance
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

A Screened Benchmark Dataset and Validity Study for Organizational Memory in Agent Harnesses

Mehul Srivastava
Zenodo (CERN European Organization for Nuclear Research)
Team Dynamics and Performance
preprint

A Screened Benchmark Dataset and Validity Study for Organizational Memory in Agent Harnesses

Mehul Srivastava
preprint en

Abstract

Organizations now use agents built on different models and software environments. Shared knowledge therefore cannot live in vendor-owned model weights; it must live in a memory layer accessible to every agent. We reviewed ten published memory benchmarks and found none tests whether memory preserves knowledge with the right scope, authority, and freshness. We introduce memory-bench, a benchmark and validity study for organizational memory. It simulates three software organizations with event streams, hidden belief-state ledgers, and behavioral work tasks. Each probe has a counterfactual twin in which the tested fact changes. An instance counts only if both versions pass. Across three released seeds, 371 instances passed screening, although all fell below the prespecified threshold of 135. A memoryless worker earned pair credit on 4 of 125 instances. Probe validity varied by worker, two harnesses with identical weights agreed on only 65% of outcomes, and judge-human agreement reached κ = 0.537, below the required 0.75. We retain these failures, document the loss of the pinned worker, and provide a re-anchoring procedure. The paper reports no comparison between memory systems.

Zenodo (CERN European Organization for Nuclear Research)
Peace, Justice and strong institutions
Team Dynamics and Performance
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

A Screened Benchmark Dataset and Validity Study for Organizational Memory in Agent Harnesses — Mehul Srivastava · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS