LLM See, LLM Do: Enhancing Repository-Aware Code Generation in LLMs Through Logical Context and Usage Knowledge

Large language models (LLMs) have significantly advanced automatic code generation. Although they have demonstrated remarkable ability in writing standalone code, they still struggle with real-world repository-level coding tasks due to the lack of repository-specific knowledge. Existing studies typically employ retrieval-augmented generation (RAG) approaches to retrieve relevant context to guide LLMs’ generation. However, these approaches fail to ensure the logical relevance and the usability of retrieved results, leading to suboptimal code generation. To address this problem, we propose LogiCoder , a novel RAG approach that leverages logical relationships within code repositories to retrieve usage knowledge for LLMs. Specifically, LogiCoder first builds a dependency graph to identify candidate callee functions. By tracing their call relationships, LogiCoder retrieves their usage examples from the entire repository. Then it integrates the most contextually similar ones with semantic search results via re-ranking. These top-ranked examples are eventually incorporated into LLMs’ prompts to enhance their repository awareness during code generation. We apply LogiCoder to five widely-used LLMs and evaluate it on the DevEval benchmark. Experimental results show that on challenging tasks that require cross-file calls, LogiCoder achieves a new state-of-the-art Pass@1 of 44.58%, with relative improvements of up to 27.26% in Pass@1 and 14.71% in Recall@1 over existing baselines. Ablation studies confirm the effectiveness of each component in LogiCoder .

Authors

Institutions

Publication Details

Journal
ACM Transactions on Software Engineering and Methodology
Published
2026-10-03
DOI
https://doi.org/10.1145/3849809
Primary Topic
Software Engineering Research
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

LLM See, LLM Do: Enhancing Repository-Aware Code Generation in LLMs Through Logical Context and Usage Knowledge

Youren Chen, Tanghaoran Zhang, Yuxin Zhao, Xinjun Mao et al.
ACM Transactions on Software Engineering and Methodology
Software Engineering Research
article

LLM See, LLM Do: Enhancing Repository-Aware Code Generation in LLMs Through Logical Context and Usage Knowledge

Youren Chen, Tanghaoran Zhang, Yuxin Zhao, Xinjun Mao, Zhang Zhang, Yao Lu, Yue Yu, Kang Yang
article en

Abstract

Large language models (LLMs) have significantly advanced automatic code generation. Although they have demonstrated remarkable ability in writing standalone code, they still struggle with real-world repository-level coding tasks due to the lack of repository-specific knowledge. Existing studies typically employ retrieval-augmented generation (RAG) approaches to retrieve relevant context to guide LLMs’ generation. However, these approaches fail to ensure the logical relevance and the usability of retrieved results, leading to suboptimal code generation. To address this problem, we propose LogiCoder , a novel RAG approach that leverages logical relationships within code repositories to retrieve usage knowledge for LLMs. Specifically, LogiCoder first builds a dependency graph to identify candidate callee functions. By tracing their call relationships, LogiCoder retrieves their usage examples from the entire repository. Then it integrates the most contextually similar ones with semantic search results via re-ranking. These top-ranked examples are eventually incorporated into LLMs’ prompts to enhance their repository awareness during code generation. We apply LogiCoder to five widely-used LLMs and evaluate it on the DevEval benchmark. Experimental results show that on challenging tasks that require cross-file calls, LogiCoder achieves a new state-of-the-art Pass@1 of 44.58%, with relative improvements of up to 27.26% in Pass@1 and 14.71% in Recall@1 over existing baselines. Ablation studies confirm the effectiveness of each component in LogiCoder .

ACM Transactions on Software Engineering and Methodology
National University of Defense Technology (CN), Peng Cheng Laboratory (CN)
Openalex Percentile: Top 5%
Software Engineering Research
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.