Language Asymmetry in Multilingual Clinical AI: A Benchmark Study of LLM Performance in ICU Decision Support Across English, Slovak, and Ukrainian

Background: Most clinical LLM benchmarks are conducted in English, on cloud infrastructure, and under conditions that bear little resemblance to real hospital environments. How these models perform in lower-resource languages — particularly in high-stakes ICU settings — remains largely unknown. Objective: To evaluate the clinical reasoning performance of a locally-deployed, RAG-augmented LLM pipeline across three languages — English (EN), Slovak (SK), and Ukrainian (UK) — using standardised synthetic ICU cases, and to identify the primary determinant of performance differences between languages. Methods: We developed GALATEA III, a three-agent clinical decision support system (Clinician → Expert → Judge) running on consumer hardware (2× NVIDIA RTX 3060 12 GB, 192 GB RAM). We benchmarked 600 synthetic ICU case evaluations (100 cases × 6 domains × 1 language per run, total 1,800 evaluations across three languages) through four optimisation phases, ending with migration from Gemma 3 12B to Gemma 4 26B. Synthetic cases were generated using GPT-4o and reviewed by the author prior to use. See Supplementary Material for full case structure and evaluation metric definitions. Results: The Gemma 3 baseline showed a clear language asymmetry: EN 40.0% / UK 48.4% / SK 18.6% APPROVE rate. Iterative optimisation reduced but did not eliminate this gap. Migration to Gemma 4 near-eliminated the asymmetry across all domains (EN 99.7% / SK 97.8% / UK 98.5%). Qualitative analysis revealed distinct reasoning styles by language: analytical-argumentative in English, procedural-directive in Slovak, and narrative-empathic in Ukrainian. Conclusions: LLM clinical performance correlates with training corpus size, not linguistic relatedness. Slovak and Ukrainian reach English-level safety with sufficient model capacity. These findings call for targeted medical corpus development for low-resource European languages.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-06-16
DOI
https://doi.org/10.5281/zenodo.20723495
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Language Asymmetry in Multilingual Clinical AI: A Benchmark Study of LLM Performance in ICU Decision Support Across English, Slovak, and Ukrainian

Taras Shlyakhta
Zenodo (CERN European Organization for Nuclear Research)
Artificial Intelligence in Healthcare and Education
article

Language Asymmetry in Multilingual Clinical AI: A Benchmark Study of LLM Performance in ICU Decision Support Across English, Slovak, and Ukrainian

Taras Shlyakhta
article en

Abstract

Background: Most clinical LLM benchmarks are conducted in English, on cloud infrastructure, and under conditions that bear little resemblance to real hospital environments. How these models perform in lower-resource languages — particularly in high-stakes ICU settings — remains largely unknown. Objective: To evaluate the clinical reasoning performance of a locally-deployed, RAG-augmented LLM pipeline across three languages — English (EN), Slovak (SK), and Ukrainian (UK) — using standardised synthetic ICU cases, and to identify the primary determinant of performance differences between languages. Methods: We developed GALATEA III, a three-agent clinical decision support system (Clinician → Expert → Judge) running on consumer hardware (2× NVIDIA RTX 3060 12 GB, 192 GB RAM). We benchmarked 600 synthetic ICU case evaluations (100 cases × 6 domains × 1 language per run, total 1,800 evaluations across three languages) through four optimisation phases, ending with migration from Gemma 3 12B to Gemma 4 26B. Synthetic cases were generated using GPT-4o and reviewed by the author prior to use. See Supplementary Material for full case structure and evaluation metric definitions. Results: The Gemma 3 baseline showed a clear language asymmetry: EN 40.0% / UK 48.4% / SK 18.6% APPROVE rate. Iterative optimisation reduced but did not eliminate this gap. Migration to Gemma 4 near-eliminated the asymmetry across all domains (EN 99.7% / SK 97.8% / UK 98.5%). Qualitative analysis revealed distinct reasoning styles by language: analytical-argumentative in English, procedural-directive in Slovak, and narrative-empathic in Ukrainian. Conclusions: LLM clinical performance correlates with training corpus size, not linguistic relatedness. Slovak and Ukrainian reach English-level safety with sufficient model capacity. These findings call for targeted medical corpus development for low-resource European languages.

Zenodo (CERN European Organization for Nuclear Research)
Uzhhorod National University (UA)
Industry, innovation and infrastructure
Openalex Percentile: Top 9%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.