Constructing a Benchmark When Every Component Is a Language Model

Synthetic benchmarks for agents often use language models to generate worlds, write tasks, perform evaluations, and judge outputs. Each dependency can change what the benchmark measures. This paper presents the engineering design of memory-bench, a benchmark for organizational memory, and measures the model dependencies that remain in its pipeline.The design keeps truth outside the models. A hidden ledger of typed, tiered, timestamped facts defines each principal’s expected belief state. Models may write event prose, author tasks and criteria, perform the work, and judge outputs, but they do not decide what is true. Counterfactual twins, mechanical ledger checks, and memoryless and fact-injected workers screen items before release.The study finds that item validity depends on the worker, agent harnesses can disagree even with identical model weights, and judge errors vary by criterion type. A multi-model judge panel showed high inter-rater agreement but did not improve validity. The pinned worker was deprecated during the study, making re-anchoring necessary. We release the construction method, quality checks, artifacts, and re-anchoring procedure. This paper reports no comparison between memory systems.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-19
DOI
https://doi.org/10.5281/zenodo.22838602
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Constructing a Benchmark When Every Component Is a Language Model

Mehul Srivastava
Zenodo (CERN European Organization for Nuclear Research)
Artificial Intelligence in Healthcare and Education
preprint

Constructing a Benchmark When Every Component Is a Language Model

Mehul Srivastava
preprint en

Abstract

Synthetic benchmarks for agents often use language models to generate worlds, write tasks, perform evaluations, and judge outputs. Each dependency can change what the benchmark measures. This paper presents the engineering design of memory-bench, a benchmark for organizational memory, and measures the model dependencies that remain in its pipeline.The design keeps truth outside the models. A hidden ledger of typed, tiered, timestamped facts defines each principal’s expected belief state. Models may write event prose, author tasks and criteria, perform the work, and judge outputs, but they do not decide what is true. Counterfactual twins, mechanical ledger checks, and memoryless and fact-injected workers screen items before release.The study finds that item validity depends on the worker, agent harnesses can disagree even with identical model weights, and judge errors vary by criterion type. A multi-model judge panel showed high inter-rater agreement but did not improve validity. The pinned worker was deprecated during the study, making re-anchoring necessary. We release the construction method, quality checks, artifacts, and re-anchoring procedure. This paper reports no comparison between memory systems.

Zenodo (CERN European Organization for Nuclear Research)
Quality Education
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Constructing a Benchmark When Every Component Is a Language Model — Mehul Srivastava · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS