MOCA: A Hierarchical Semantic-Enhanced Code Edit Framework for Multilingual Code Co-Evolution
Multilingual code co-evolution aims to propagate code edits across programming languages while preserving functional consistency, yet the task remains challenging in real-world repositories. Our motivating examples reveal two key challenges: (1) Existing models fail to infer the rationale behind source edits and may omit or misidentify the corresponding edits in the target code; and (2) Models do not capture repository-level semantics such as API usages and cross-file dependencies required for correct adaptations. To address these limitations, we introduce MOCA, a hierarchical semantic-enhanced framework that operationalizes two complementary strategies: (1) edit-wise intent interpretation, which explicates the intent of each source edit and guides its cross-lingual mapping, and (2) project-wise context integration, which injects repository-level semantics to ensure that propagated edits adhere to project constraints. MOCA implements these strategies through four coordinated steps: a Summarizer that provides multi-perspective explanations of each source edit, a Locator that identifies corresponding edit regions in the target code, a Retriever that supplies and filters repository context, and a Modifier that integrates all reasoning signals to synthesize the final edits. Experiments on eight real-world Java–C# projects demonstrate that MOCA achieves substantial improvements under both unseen-project ( \(S_{proj}\) ) and time-split seen-project ( \(S_{time}\) ) settings, reaching 73.06%/71.84% exact match for C#→Java/Java→C# under \(S_{proj}\) and 85.99%/79.36% under \(S_{time}\) . Compared with the state-of-the-art fine-tuning-based method Codeditor, MOCA improves exact match by 11.69–32.04%; compared with prompt-engineering-based baselines, it further yields 5.10–6.54% under \(S_{proj}\) and 5.63–5.81% under \(S_{time}\) . Our ablation studies confirm the contribution of all designed agents, while generalization experiments show that MOCA consistently outperforms each model's strongest prompt baseline by 3.56–15.63%. Qualitative analyses further indicate that MOCA handles a broader range of multi-hunk and semantically complex edits. Overall, MOCA provides a practical paradigm for multilingual code co-evolution by jointly strengthening edit-wise semantics and project-wise grounding.
Authors
- Zeyu Sun (ORCID: https://orcid.org/0000-0002-9990-9120)
- Dan Hao (ORCID: https://orcid.org/0000-0001-8295-303X)
- Xin Yin
- Yizhou Chen (ORCID: https://orcid.org/0000-0003-1821-3170)
- Zhihao Gong
- Guoqing Wang
- Qingyuan Liang
- Jie Zhang
Institutions
- King's College London (GB)
- Chinese Academy of Sciences (CN)
- Peking University (CN)
- Institute of Software (CN)
- Zhejiang University (CN)
Publication Details
- Journal
- ACM Transactions on Software Engineering and Methodology
- Published
- 2026-10-08
- DOI
- https://doi.org/10.1145/3849703
- Primary Topic
- Software Engineering Research
- Type
- article
- Field-Weighted Citation Impact
- 0.00