LinearBench: A Framework for Measuring Convergence, Stability, and Feedback Dependence in Iterative Product Construction

LinearBench is a framework for evaluating language models on multi-round build-and-refine tasks, in which an artifact is revised under feedback and judged against a hidden, weighted acceptance specification. A deterministic user simulator supplies feedback at four levels of informativeness, from generic self-review to machine-readable diagnostics, which separates a model's capacity for self-correction from its dependence on external guidance. The paper defines trajectory-level metrics, including stable turns-to-acceptance under right censoring, a regression and churn decomposition, a headroom-weighted feedback efficiency, a parameter-free persistence-based Convergence Score, and a Self-Sufficiency Ratio, and proves that summaries computed from the quality sequence alone cannot detect regressions. A synthetic study with parameterized model profiles illustrates cases in which conventional and proposed summaries rank the same profiles differently. This is a measurement proposal that reports no evaluation of existing language models; it specifies a pilot design, falsifiable hypotheses, and open problems, and invites collaboration on task construction, verification, and human validation of the simulator.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-03
DOI
https://doi.org/10.5281/zenodo.23124587
Primary Topic
Topic Modeling
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

LinearBench: A Framework for Measuring Convergence, Stability, and Feedback Dependence in Iterative Product Construction

Momen Ghazouani
Zenodo (CERN European Organization for Nuclear Research)
Topic Modeling
preprint

LinearBench: A Framework for Measuring Convergence, Stability, and Feedback Dependence in Iterative Product Construction

Momen Ghazouani
preprint en

Abstract

LinearBench is a framework for evaluating language models on multi-round build-and-refine tasks, in which an artifact is revised under feedback and judged against a hidden, weighted acceptance specification. A deterministic user simulator supplies feedback at four levels of informativeness, from generic self-review to machine-readable diagnostics, which separates a model's capacity for self-correction from its dependence on external guidance. The paper defines trajectory-level metrics, including stable turns-to-acceptance under right censoring, a regression and churn decomposition, a headroom-weighted feedback efficiency, a parameter-free persistence-based Convergence Score, and a Self-Sufficiency Ratio, and proves that summaries computed from the quality sequence alone cannot detect regressions. A synthetic study with parameterized model profiles illustrates cases in which conventional and proposed summaries rank the same profiles differently. This is a measurement proposal that reports no evaluation of existing language models; it specifies a pilot design, falsifiable hypotheses, and open problems, and invites collaboration on task construction, verification, and human validation of the simulator.

Zenodo (CERN European Organization for Nuclear Research)
Topic Modeling
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.