Constructing a Benchmark When Every Component Is a Language Model
Synthetic benchmarks for agents often use language models to generate worlds, write tasks, perform evaluations, and judge outputs. Each dependency can change what the benchmark measures. This paper presents the engineering design of memory-bench, a benchmark for organizational memory, and measures the model dependencies that remain in its pipeline.The design keeps truth outside the models. A hidden ledger of typed, tiered, timestamped facts defines each principal’s expected belief state. Models may write event prose, author tasks and criteria, perform the work, and judge outputs, but they do not decide what is true. Counterfactual twins, mechanical ledger checks, and memoryless and fact-injected workers screen items before release.The study finds that item validity depends on the worker, agent harnesses can disagree even with identical model weights, and judge errors vary by criterion type. A multi-model judge panel showed high inter-rater agreement but did not improve validity. The pinned worker was deprecated during the study, making re-anchoring necessary. We release the construction method, quality checks, artifacts, and re-anchoring procedure. This paper reports no comparison between memory systems.
Authors
- Mehul Srivastava (ORCID: https://orcid.org/0009-0008-1031-304X)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-19
- DOI
- https://doi.org/10.5281/zenodo.22838603
- Primary Topic
- Artificial Intelligence in Healthcare and Education
- Type
- preprint