Large Language Models for Code Generation from Multilingual Prompts: A Curated Benchmark and a Study on Code Quality
Large Language Models (LLMs) perform differently on identical tasks when prompted in different languages—a phenomenon known as language bias. Although well-documented for general text generation, the extent to which language bias affects code generation quality and programming conventions remains largely unexplored. To address this gap, we investigate how the natural language used to describe programming tasks influences the quality of source code generated by three prominent LLMs: GPT-4o mini , DeepSeek, and Claude. Our study includes 460 coding tasks spanning Python (230 tasks) and Java (230 tasks). We translate and manually curate the original English prompts into four diverse languages: Chinese, Hindi, Spanish, and Italian, ensuring linguistic accuracy while preserving technical meaning. We evaluate the resulting code across multiple dimensions: functional correctness through test passage rates, structural quality via established code metrics, potential issues identified by static analysis tools, and lexical characteristics, including the natural language used in identifiers and comments. Results indicate that (i) source code generated from the English queries is not necessarily better in terms of passed test and quality metrics, (ii) the quality for different languages varies depending on the programming language and LLM being used, and (iii) the generated code tends to contain mixes of comments and literals written in English and the prompt language.
Authors
- Massimiliano Di Penta (ORCID: https://orcid.org/0000-0002-0340-9747)
- Camilo Escobar‐Velásquez (ORCID: https://orcid.org/0000-0001-8414-9301)
- Alessandro Midolo (ORCID: https://orcid.org/0000-0002-9575-8054)
- Weiyuan Ding (ORCID: https://orcid.org/0000-0003-3232-0693)
- Mario Linares‐Vásquez (ORCID: https://orcid.org/0000-0003-0161-2888)
- Antonio Mastropaolo (ORCID: https://orcid.org/0000-0002-7965-7712)
- Saima Afrin (ORCID: https://orcid.org/0009-0008-4106-6838)
- Bowen Xu
Institutions
- North Carolina State University (US)
- Universidad de Los Andes (CO)
- William & Mary (US)
- University of Sannio (IT)
- University of Catania (IT)
Publication Details
- Journal
- ACM Transactions on Software Engineering and Methodology
- Published
- 2026-10-06
- DOI
- https://doi.org/10.1145/3845994
- Primary Topic
- Software Engineering Research
- Type
- article
- Field-Weighted Citation Impact
- 0.00