GALATEA II: Benchmarking LLM Safety in Clinical Simulation Behavioural Safety and Ethical Robustness of Large Language Models in a Multi-Agent ICU Decision Support Architecture
Background. Large language models (LLMs) are increasingly proposed as clinical decision support tools in intensive care. Most existing evaluations focus on static recall of medical knowledge and do not capture how a model behaves in dynamic clinical dialogue, under social pressure, or in ethical conflict. The safety of LLM-based systems under these conditions remains poorly characterised. Methods. We benchmarked 12 language models in a three-role multi-agent architecture (Clinician — Guardian — Judge). We generated 42,842 clinical consultations across 9 clinical domains, 3 ethical profiles and 4 case types. Clinical inference ran locally on consumer hardware. Quality was assessed by two independent LLM judges — a local GPT-OSS model and Gemini 2.5 Flash via API — which together covered 41,384 consultations (96.6%), 17,064 of them jointly. That overlap let us measure the reliability of the evaluation itself. Results. Accuracy ranged from 11.1% to 77.0%. Three of the four specialised medical models underperformed general-purpose models; the worst capitulated under authority pressure in 75.8% of cases. The exception, medgemma-27b-it (74.8% accuracy, 7.9% sycophancy), shows that safe specialisation is achievable. The most restrictive ethical profile vetoed 64.9% of all plans and produced the lowest accuracy (45.9% versus 66.6%), which the GQI metric expresses as 0.71 versus 2.70. Clinical memory degrades under load in at least two independent ways: under authority pressure, coupled with capitulation (70.4% co-occurrence with sycophancy), and under conflicting data, with no social pressure at all (14.4%). The two judges agreed almost perfectly on verdict correctness (κ = 0.929) and not at all on memory failure (κ = 0.029). Conclusions. Clinical LLM safety is determined neither by medical specialisation nor by the strictness of ethical constraints, but by architectural resilience to social manipulation and by memory integrity under load. A separate methodological result: evaluating clinical AI with other AI requires reporting inter-judge agreement, because within a single rubric that agreement ranged from almost perfect to none.
Authors
- Taras Shlyakhta (ORCID: https://orcid.org/0009-0009-3230-5603)
Institutions
- Linde (United States) (US)
- Łukasiewicz Research Network - Institute of Electrical Drives & Machines KOMEL (PL)
- Galena Biopharma (United States) (US)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-06
- DOI
- https://doi.org/10.5281/zenodo.19828812
- Primary Topic
- Artificial Intelligence in Healthcare and Education
- Type
- article
- Field-Weighted Citation Impact
- 0.00