Enhancing LLM-Based Bug Reproduction via Code Entity Retrieval and Test Case Repair

Automated bug reproduction from bug reports is a critical yet challenging step in software debugging. While LLM-based bug reproduction shows promise, its effectiveness is often hampered by insufficient contextual awareness of the relevant codebase and a tendency to produce invalid test cases. To address these limitations, we propose a novel approach, called LTER, that enhances LLM-based bug reproduction through fine-grained code entity retrieval and a feedback-driven dynamic repair loop. LTER first identifies specific code entities within bug reports to automatically extract precise contexts, including class definitions, constructors, and method logic. The extracted contexts are then used to guide the LLM in generating reproduced test cases. To further ensure executability, LTER employs an iterative repair mechanism to resolve complex dependencies. Specifically, upon injecting a generated test case into the project, if a compilation failure occurs, the framework forwards the error messages to the LLM for an initial repair. Should this initial repair fail, it empowers the LLM to analyze diagnostic messages to recognize missing context and retrieve indispensable dependencies, subsequently regenerating the test case with the supplemented data. Finally, LTER employs a hybrid cascade ranking strategy to accurately select the most effective reproduction test case from the generated candidates. The experimental results on the widely-used Defects4J benchmark show that LTER substantially outperforms the best performance in automated bug reproduction, increasing the reproduction success rate to 46.2% with successfully identifying a valid reproduction test as the top candidate in 38.1% of the cases. Furthermore, LTER demonstrates strong generalization capability, delivering robust performance on the GHRB dataset containing recent bugs previously unseen by the LLM.

Authors

Institutions

Publication Details

Journal
Proceedings of the ACM on software engineering.
Published
2026-10-01
DOI
https://doi.org/10.1145/3832247
Primary Topic
Software Testing and Debugging Techniques
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Enhancing LLM-Based Bug Reproduction via Code Entity Retrieval and Test Case Repair

Hui Liu, Yanjie Jiang, Yuxia Zhang, Hao Ding
Proceedings of the ACM on software engineering.
Software Testing and Debugging Techniques
article

Enhancing LLM-Based Bug Reproduction via Code Entity Retrieval and Test Case Repair

Hui Liu, Yanjie Jiang, Yuxia Zhang, Hao Ding
article en

Abstract

Automated bug reproduction from bug reports is a critical yet challenging step in software debugging. While LLM-based bug reproduction shows promise, its effectiveness is often hampered by insufficient contextual awareness of the relevant codebase and a tendency to produce invalid test cases. To address these limitations, we propose a novel approach, called LTER, that enhances LLM-based bug reproduction through fine-grained code entity retrieval and a feedback-driven dynamic repair loop. LTER first identifies specific code entities within bug reports to automatically extract precise contexts, including class definitions, constructors, and method logic. The extracted contexts are then used to guide the LLM in generating reproduced test cases. To further ensure executability, LTER employs an iterative repair mechanism to resolve complex dependencies. Specifically, upon injecting a generated test case into the project, if a compilation failure occurs, the framework forwards the error messages to the LLM for an initial repair. Should this initial repair fail, it empowers the LLM to analyze diagnostic messages to recognize missing context and retrieve indispensable dependencies, subsequently regenerating the test case with the supplemented data. Finally, LTER employs a hybrid cascade ranking strategy to accurately select the most effective reproduction test case from the generated candidates. The experimental results on the widely-used Defects4J benchmark show that LTER substantially outperforms the best performance in automated bug reproduction, increasing the reproduction success rate to 46.2% with successfully identifying a valid reproduction test as the top candidate in 38.1% of the cases. Furthermore, LTER demonstrates strong generalization capability, delivering robust performance on the GHRB dataset containing recent bugs previously unseen by the LLM.

Proceedings of the ACM on software engineering.Vol. 3(ISSTA)
Beijing Institute of Technology (CN), Tianjin University (CN)
Openalex Percentile: Top 7%
Software Testing and Debugging Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.