GPT-4o Translation of Pediatric Discharge Instructions Across Languages
BACKGROUND AND OBJECTIVES: Patients and caregivers who use languages other than English in health care settings rarely receive language-concordant discharge instructions. Large language models show promise for improving access to translations but must be validated to ensure their output is safe for clinical use. This study measured the quality of GPT-4o medical translations to Portuguese, Haitian Creole, and Arabic. METHODS: Previously translated personalized discharge instructions (performed by professional human translators) were extracted from inpatient admissions at a large US pediatric academic medical center. The discharge instructions were translated to Portuguese, Haitian Creole, and Arabic by GPT-4o. Medical translators then evaluated the translations using both the Multidimensional Quality Metrics (MQM) framework and a preference scale. RESULTS: Translations of 14 to 20 discharge instructions for each of the 3 languages were evaluated. Noninferiority testing showed no significant difference in translation quality between GPT-4o and human translations for Portuguese and Arabic, but worse performance by GPT-4o than human translations for Haitian Creole (mean MQM score 90.5 ± SD 5.5 vs 97.0 ± SD 2.1). Linguists preferred the human translation over GPT-4o in 69% (SE = 7%) of the Haitian Creole cases. CONCLUSIONS: In this cross-sectional study, GPT-4o translation performance of pediatric discharge instructions was noninferior for Portuguese and Arabic but was worse in Haitian Creole. Human validation of medical translations remains necessary for all languages, but large language models can be used to draft high-quality translations in languages with proven high performance, while resources are shifted toward more intensive post-editing of translations in languages with lower performance.
Authors
- Ryan Brewster (ORCID: https://orcid.org/0000-0003-0051-2623)
- Mondira Ray (ORCID: https://orcid.org/0000-0003-1248-6610)
- Benjamin M. Rader (ORCID: https://orcid.org/0000-0002-6095-0193)
- Joss Moorkens (ORCID: https://orcid.org/0000-0003-0766-0071)
- Daniel J. Kats (ORCID: https://orcid.org/0000-0001-7591-6222)
- Jonathan D. Hron (ORCID: https://orcid.org/0000-0003-2198-6174)
- John S. Brownstein (ORCID: https://orcid.org/0000-0001-8568-5317)
- Dinesh Rai
- Alisa Khan
Institutions
- Boston Children's Hospital (US)
- Beth Israel Deaconess Medical Center (US)
- Harvard University (US)
- Dublin City University (IE)
Publication Details
- Journal
- Hospital Pediatrics
- Published
- 2026-10-08
- DOI
- https://doi.org/10.1542/hpeds.2026-009476
- Primary Topic
- Interpreting and Communication in Healthcare
- Type
- article
- Field-Weighted Citation Impact
- 0.00