Large Language Models for Code Generation from Multilingual Prompts: A Curated Benchmark and a Study on Code Quality

Large Language Models (LLMs) perform differently on identical tasks when prompted in different languages—a phenomenon known as language bias. Although well-documented for general text generation, the extent to which language bias affects code generation quality and programming conventions remains largely unexplored. To address this gap, we investigate how the natural language used to describe programming tasks influences the quality of source code generated by three prominent LLMs: GPT-4o mini , DeepSeek, and Claude. Our study includes 460 coding tasks spanning Python (230 tasks) and Java (230 tasks). We translate and manually curate the original English prompts into four diverse languages: Chinese, Hindi, Spanish, and Italian, ensuring linguistic accuracy while preserving technical meaning. We evaluate the resulting code across multiple dimensions: functional correctness through test passage rates, structural quality via established code metrics, potential issues identified by static analysis tools, and lexical characteristics, including the natural language used in identifiers and comments. Results indicate that (i) source code generated from the English queries is not necessarily better in terms of passed test and quality metrics, (ii) the quality for different languages varies depending on the programming language and LLM being used, and (iii) the generated code tends to contain mixes of comments and literals written in English and the prompt language.

Authors

Institutions

Publication Details

Journal
ACM Transactions on Software Engineering and Methodology
Published
2026-10-06
DOI
https://doi.org/10.1145/3845994
Primary Topic
Software Engineering Research
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Large Language Models for Code Generation from Multilingual Prompts: A Curated Benchmark and a Study on Code Quality

Massimiliano Di Penta, Camilo Escobar‐Velásquez, Alessandro Midolo, Weiyuan Ding et al.
ACM Transactions on Software Engineering and Methodology
Software Engineering Research
article

Large Language Models for Code Generation from Multilingual Prompts: A Curated Benchmark and a Study on Code Quality

Massimiliano Di Penta, Camilo Escobar‐Velásquez, Alessandro Midolo, Weiyuan Ding, Mario Linares‐Vásquez, Antonio Mastropaolo, Saima Afrin, Bowen Xu
article en

Abstract

Large Language Models (LLMs) perform differently on identical tasks when prompted in different languages—a phenomenon known as language bias. Although well-documented for general text generation, the extent to which language bias affects code generation quality and programming conventions remains largely unexplored. To address this gap, we investigate how the natural language used to describe programming tasks influences the quality of source code generated by three prominent LLMs: GPT-4o mini , DeepSeek, and Claude. Our study includes 460 coding tasks spanning Python (230 tasks) and Java (230 tasks). We translate and manually curate the original English prompts into four diverse languages: Chinese, Hindi, Spanish, and Italian, ensuring linguistic accuracy while preserving technical meaning. We evaluate the resulting code across multiple dimensions: functional correctness through test passage rates, structural quality via established code metrics, potential issues identified by static analysis tools, and lexical characteristics, including the natural language used in identifiers and comments. Results indicate that (i) source code generated from the English queries is not necessarily better in terms of passed test and quality metrics, (ii) the quality for different languages varies depending on the programming language and LLM being used, and (iii) the generated code tends to contain mixes of comments and literals written in English and the prompt language.

ACM Transactions on Software Engineering and Methodology
North Carolina State University (US), Universidad de Los Andes (CO), William & Mary (US), University of Sannio (IT), University of Catania (IT)
Openalex Percentile: Top 41%
Software Engineering Research
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.