GPT-4o Translation of Pediatric Discharge Instructions Across Languages

BACKGROUND AND OBJECTIVES: Patients and caregivers who use languages other than English in health care settings rarely receive language-concordant discharge instructions. Large language models show promise for improving access to translations but must be validated to ensure their output is safe for clinical use. This study measured the quality of GPT-4o medical translations to Portuguese, Haitian Creole, and Arabic. METHODS: Previously translated personalized discharge instructions (performed by professional human translators) were extracted from inpatient admissions at a large US pediatric academic medical center. The discharge instructions were translated to Portuguese, Haitian Creole, and Arabic by GPT-4o. Medical translators then evaluated the translations using both the Multidimensional Quality Metrics (MQM) framework and a preference scale. RESULTS: Translations of 14 to 20 discharge instructions for each of the 3 languages were evaluated. Noninferiority testing showed no significant difference in translation quality between GPT-4o and human translations for Portuguese and Arabic, but worse performance by GPT-4o than human translations for Haitian Creole (mean MQM score 90.5 ± SD 5.5 vs 97.0 ± SD 2.1). Linguists preferred the human translation over GPT-4o in 69% (SE = 7%) of the Haitian Creole cases. CONCLUSIONS: In this cross-sectional study, GPT-4o translation performance of pediatric discharge instructions was noninferior for Portuguese and Arabic but was worse in Haitian Creole. Human validation of medical translations remains necessary for all languages, but large language models can be used to draft high-quality translations in languages with proven high performance, while resources are shifted toward more intensive post-editing of translations in languages with lower performance.

Authors

Institutions

Publication Details

Journal
Hospital Pediatrics
Published
2026-10-08
DOI
https://doi.org/10.1542/hpeds.2026-009476
Primary Topic
Interpreting and Communication in Healthcare
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

GPT-4o Translation of Pediatric Discharge Instructions Across Languages

Ryan Brewster, Mondira Ray, Benjamin M. Rader, Joss Moorkens et al.
Hospital Pediatrics
Interpreting and Communication in Healthcare
article

GPT-4o Translation of Pediatric Discharge Instructions Across Languages

Ryan Brewster, Mondira Ray, Benjamin M. Rader, Joss Moorkens, Daniel J. Kats, Jonathan D. Hron, John S. Brownstein, Dinesh Rai, Alisa Khan
article en

Abstract

BACKGROUND AND OBJECTIVES: Patients and caregivers who use languages other than English in health care settings rarely receive language-concordant discharge instructions. Large language models show promise for improving access to translations but must be validated to ensure their output is safe for clinical use. This study measured the quality of GPT-4o medical translations to Portuguese, Haitian Creole, and Arabic. METHODS: Previously translated personalized discharge instructions (performed by professional human translators) were extracted from inpatient admissions at a large US pediatric academic medical center. The discharge instructions were translated to Portuguese, Haitian Creole, and Arabic by GPT-4o. Medical translators then evaluated the translations using both the Multidimensional Quality Metrics (MQM) framework and a preference scale. RESULTS: Translations of 14 to 20 discharge instructions for each of the 3 languages were evaluated. Noninferiority testing showed no significant difference in translation quality between GPT-4o and human translations for Portuguese and Arabic, but worse performance by GPT-4o than human translations for Haitian Creole (mean MQM score 90.5 ± SD 5.5 vs 97.0 ± SD 2.1). Linguists preferred the human translation over GPT-4o in 69% (SE = 7%) of the Haitian Creole cases. CONCLUSIONS: In this cross-sectional study, GPT-4o translation performance of pediatric discharge instructions was noninferior for Portuguese and Arabic but was worse in Haitian Creole. Human validation of medical translations remains necessary for all languages, but large language models can be used to draft high-quality translations in languages with proven high performance, while resources are shifted toward more intensive post-editing of translations in languages with lower performance.

Hospital Pediatrics
Boston Children's Hospital (US), Beth Israel Deaconess Medical Center (US), Harvard University (US), Dublin City University (IE)
Openalex Percentile: Top 8%
Interpreting and Communication in Healthcare
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.