GALATEA II: Benchmarking LLM Safety in Clinical Simulation Behavioural Safety and Ethical Robustness of Large Language Models in a Multi-Agent ICU Decision Support Architecture

Background. Large language models (LLMs) are increasingly proposed as clinical decision support tools in intensive care. Most existing evaluations focus on static recall of medical knowledge and do not capture how a model behaves in dynamic clinical dialogue, under social pressure, or in ethical conflict. The safety of LLM-based systems under these conditions remains poorly characterised. Methods. We benchmarked 12 language models in a three-role multi-agent architecture (Clinician — Guardian — Judge). We generated 42,842 clinical consultations across 9 clinical domains, 3 ethical profiles and 4 case types. Clinical inference ran locally on consumer hardware. Quality was assessed by two independent LLM judges — a local GPT-OSS model and Gemini 2.5 Flash via API — which together covered 41,384 consultations (96.6%), 17,064 of them jointly. That overlap let us measure the reliability of the evaluation itself. Results. Accuracy ranged from 11.1% to 77.0%. Three of the four specialised medical models underperformed general-purpose models; the worst capitulated under authority pressure in 75.8% of cases. The exception, medgemma-27b-it (74.8% accuracy, 7.9% sycophancy), shows that safe specialisation is achievable. The most restrictive ethical profile vetoed 64.9% of all plans and produced the lowest accuracy (45.9% versus 66.6%), which the GQI metric expresses as 0.71 versus 2.70. Clinical memory degrades under load in at least two independent ways: under authority pressure, coupled with capitulation (70.4% co-occurrence with sycophancy), and under conflicting data, with no social pressure at all (14.4%). The two judges agreed almost perfectly on verdict correctness (κ = 0.929) and not at all on memory failure (κ = 0.029). Conclusions. Clinical LLM safety is determined neither by medical specialisation nor by the strictness of ethical constraints, but by architectural resilience to social manipulation and by memory integrity under load. A separate methodological result: evaluating clinical AI with other AI requires reporting inter-judge agreement, because within a single rubric that agreement ranged from almost perfect to none.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-06
DOI
https://doi.org/10.5281/zenodo.19828812
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

GALATEA II: Benchmarking LLM Safety in Clinical Simulation Behavioural Safety and Ethical Robustness of Large Language Models in a Multi-Agent ICU Decision Support Architecture

Taras Shlyakhta
Zenodo (CERN European Organization for Nuclear Research)
Artificial Intelligence in Healthcare and Education
article

GALATEA II: Benchmarking LLM Safety in Clinical Simulation Behavioural Safety and Ethical Robustness of Large Language Models in a Multi-Agent ICU Decision Support Architecture

Taras Shlyakhta
article en

Abstract

Background. Large language models (LLMs) are increasingly proposed as clinical decision support tools in intensive care. Most existing evaluations focus on static recall of medical knowledge and do not capture how a model behaves in dynamic clinical dialogue, under social pressure, or in ethical conflict. The safety of LLM-based systems under these conditions remains poorly characterised. Methods. We benchmarked 12 language models in a three-role multi-agent architecture (Clinician — Guardian — Judge). We generated 42,842 clinical consultations across 9 clinical domains, 3 ethical profiles and 4 case types. Clinical inference ran locally on consumer hardware. Quality was assessed by two independent LLM judges — a local GPT-OSS model and Gemini 2.5 Flash via API — which together covered 41,384 consultations (96.6%), 17,064 of them jointly. That overlap let us measure the reliability of the evaluation itself. Results. Accuracy ranged from 11.1% to 77.0%. Three of the four specialised medical models underperformed general-purpose models; the worst capitulated under authority pressure in 75.8% of cases. The exception, medgemma-27b-it (74.8% accuracy, 7.9% sycophancy), shows that safe specialisation is achievable. The most restrictive ethical profile vetoed 64.9% of all plans and produced the lowest accuracy (45.9% versus 66.6%), which the GQI metric expresses as 0.71 versus 2.70. Clinical memory degrades under load in at least two independent ways: under authority pressure, coupled with capitulation (70.4% co-occurrence with sycophancy), and under conflicting data, with no social pressure at all (14.4%). The two judges agreed almost perfectly on verdict correctness (κ = 0.929) and not at all on memory failure (κ = 0.029). Conclusions. Clinical LLM safety is determined neither by medical specialisation nor by the strictness of ethical constraints, but by architectural resilience to social manipulation and by memory integrity under load. A separate methodological result: evaluating clinical AI with other AI requires reporting inter-judge agreement, because within a single rubric that agreement ranged from almost perfect to none.

Zenodo (CERN European Organization for Nuclear Research)
Linde (United States) (US), Łukasiewicz Research Network - Institute of Electrical Drives & Machines KOMEL (PL), Galena Biopharma (United States) (US)
Peace, Justice and strong institutions
Openalex Percentile: Top 69%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.