Baseline Output Length for Prioritizing Visible-Test Regressions Under a Task-Count Budget: Controlled Model and Retrieval-Context Changes in LLM-Enabled Code Generation

Non-code AI dependencies of an LLM-enabled code-generation pipeline, such as the model checkpoint and the retrieved context, can change without any application-code commit. When only a fixed number of task re-executions can be afforded, the operational question is which tasks to re-run first. We study a training-free ordering that uses only pre-update baseline logs: the mean raw-response character length. A controlled derivation block over 150 MBPP tasks evaluates two retrieval-context interventions and one model swap; a separate frozen task-disjoint block applies target-answer-excluded contexts and quantized execution, so that task population, context construction and execution regime change together. Baseline output length improved prioritization relative to random ordering under the two controlled retrieval-context interventions, whereas the derivation-block model-swap estimate was favorable but imprecise. In that compound-shift block the prespecified criterion was not met (APFD 0.597 against a 0.60 threshold); the corresponding AUROC was 0.640 over the 66.7 percent of panel tasks that were eligible. A later bridging experiment with only three regression events did not isolate the contribution of the task population. All labels are visible-test outcomes, and the budget is a count of task re-executions rather than token, latency, or total operational cost.

Authors

Institutions

Publication Details

Journal
Electronics
Published
2026-09-21
DOI
https://doi.org/10.3390/electronics15184334
Primary Topic
Software Engineering Research
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Baseline Output Length for Prioritizing Visible-Test Regressions Under a Task-Count Budget: Controlled Model and Retrieval-Context Changes in LLM-Enabled Code Generation

Geunseok Yang, Gyumin Nam
Electronics
Software Engineering Research
article

Baseline Output Length for Prioritizing Visible-Test Regressions Under a Task-Count Budget: Controlled Model and Retrieval-Context Changes in LLM-Enabled Code Generation

Geunseok Yang, Gyumin Nam
article en

Abstract

Non-code AI dependencies of an LLM-enabled code-generation pipeline, such as the model checkpoint and the retrieved context, can change without any application-code commit. When only a fixed number of task re-executions can be afforded, the operational question is which tasks to re-run first. We study a training-free ordering that uses only pre-update baseline logs: the mean raw-response character length. A controlled derivation block over 150 MBPP tasks evaluates two retrieval-context interventions and one model swap; a separate frozen task-disjoint block applies target-answer-excluded contexts and quantized execution, so that task population, context construction and execution regime change together. Baseline output length improved prioritization relative to random ordering under the two controlled retrieval-context interventions, whereas the derivation-block model-swap estimate was favorable but imprecise. In that compound-shift block the prespecified criterion was not met (APFD 0.597 against a 0.60 threshold); the corresponding AUROC was 0.640 over the 66.7 percent of panel tasks that were eligible. A later bridging experiment with only three regression events did not isolate the contribution of the task population. All labels are visible-test outcomes, and the budget is a count of task re-executions rather than token, latency, or total operational cost.

ElectronicsVol. 15(18)
Hankyong National University (KR)
Openalex Percentile: Top 4%
Software Engineering Research
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.