LinearBench: A Framework for Measuring Convergence, Stability, and Feedback Dependence in Iterative Product Construction
LinearBench is a framework for evaluating language models on multi-round build-and-refine tasks, in which an artifact is revised under feedback and judged against a hidden, weighted acceptance specification. A deterministic user simulator supplies feedback at four levels of informativeness, from generic self-review to machine-readable diagnostics, which separates a model's capacity for self-correction from its dependence on external guidance. The paper defines trajectory-level metrics, including stable turns-to-acceptance under right censoring, a regression and churn decomposition, a headroom-weighted feedback efficiency, a parameter-free persistence-based Convergence Score, and a Self-Sufficiency Ratio, and proves that summaries computed from the quality sequence alone cannot detect regressions. A synthetic study with parameterized model profiles illustrates cases in which conventional and proposed summaries rank the same profiles differently. This is a measurement proposal that reports no evaluation of existing language models; it specifies a pilot design, falsifiable hypotheses, and open problems, and invites collaboration on task construction, verification, and human validation of the simulator.
Authors
- Momen Ghazouani (ORCID: https://orcid.org/0009-0003-9484-9899)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-03
- DOI
- https://doi.org/10.5281/zenodo.23124588
- Primary Topic
- Topic Modeling
- Type
- preprint