Baseline Output Length for Prioritizing Visible-Test Regressions Under a Task-Count Budget: Controlled Model and Retrieval-Context Changes in LLM-Enabled Code Generation
Non-code AI dependencies of an LLM-enabled code-generation pipeline, such as the model checkpoint and the retrieved context, can change without any application-code commit. When only a fixed number of task re-executions can be afforded, the operational question is which tasks to re-run first. We study a training-free ordering that uses only pre-update baseline logs: the mean raw-response character length. A controlled derivation block over 150 MBPP tasks evaluates two retrieval-context interventions and one model swap; a separate frozen task-disjoint block applies target-answer-excluded contexts and quantized execution, so that task population, context construction and execution regime change together. Baseline output length improved prioritization relative to random ordering under the two controlled retrieval-context interventions, whereas the derivation-block model-swap estimate was favorable but imprecise. In that compound-shift block the prespecified criterion was not met (APFD 0.597 against a 0.60 threshold); the corresponding AUROC was 0.640 over the 66.7 percent of panel tasks that were eligible. A later bridging experiment with only three regression events did not isolate the contribution of the task population. All labels are visible-test outcomes, and the budget is a count of task re-executions rather than token, latency, or total operational cost.
Authors
- Geunseok Yang (ORCID: https://orcid.org/0000-0001-5677-5129)
- Gyumin Nam (ORCID: https://orcid.org/0009-0005-9385-7234)
Institutions
- Hankyong National University (KR)
Publication Details
- Journal
- Electronics
- Published
- 2026-09-21
- DOI
- https://doi.org/10.3390/electronics15184334
- Primary Topic
- Software Engineering Research
- Type
- article
- Field-Weighted Citation Impact
- 0.00