Benchmarking the safety of large language models for robotic health attendant control

Abstract Large language models (LLMs) are increasingly considered for deployment as the control component of robotic health attendants (RHAs), yet their safety in this context remains poorly characterized. We introduce a dataset of 270 harmful instructions spanning nine prohibited behaviour categories grounded in the American Medical Association (AMA) Principles of Medical Ethics, and use it to evaluate 72 LLMs in a simulation environment based on the RHA framework. The mean violation rate across all models was 54.4%, with more than half exceeding 50%, and violation rates varied substantially across behaviour categories, with superficially plausible instructions such as device manipulation and emergency delay proving harder to refuse than overtly destructive ones. Model size and release date were the primary determinants of safety performance among open-weight models, and proprietary models were substantially safer than open-weight counterparts (median 23.7% versus 72.8%). Medical domain fine-tuning conferred no significant overall safety benefit, and a prompt-based defence strategy produced only a modest reduction in violation rates among the least safe models, leaving violation rates that would preclude safe clinical deployment. These findings demonstrate that safety evaluation must be treated as a first-class criterion in the development and deployment of LLM-based RHAs.

Authors

Institutions

Publication Details

Journal
Royal Society Open Science
Published
2026-09-30
DOI
https://doi.org/10.1098/rsos.261022
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Benchmarking the safety of large language models for robotic health attendant control

Kazuhiro Takemoto, Mahiro Nakao
Royal Society Open Science
Artificial Intelligence in Healthcare and Education
article

Benchmarking the safety of large language models for robotic health attendant control

Kazuhiro Takemoto, Mahiro Nakao
article en

Abstract

Abstract Large language models (LLMs) are increasingly considered for deployment as the control component of robotic health attendants (RHAs), yet their safety in this context remains poorly characterized. We introduce a dataset of 270 harmful instructions spanning nine prohibited behaviour categories grounded in the American Medical Association (AMA) Principles of Medical Ethics, and use it to evaluate 72 LLMs in a simulation environment based on the RHA framework. The mean violation rate across all models was 54.4%, with more than half exceeding 50%, and violation rates varied substantially across behaviour categories, with superficially plausible instructions such as device manipulation and emergency delay proving harder to refuse than overtly destructive ones. Model size and release date were the primary determinants of safety performance among open-weight models, and proprietary models were substantially safer than open-weight counterparts (median 23.7% versus 72.8%). Medical domain fine-tuning conferred no significant overall safety benefit, and a prompt-based defence strategy produced only a modest reduction in violation rates among the least safe models, leaving violation rates that would preclude safe clinical deployment. These findings demonstrate that safety evaluation must be treated as a first-class criterion in the development and deployment of LLM-based RHAs.

Royal Society Open ScienceVol. 13(9)
Kyushu Institute of Technology (JP)
Japan Society for the Promotion of Science
Peace, Justice and strong institutions
Openalex Percentile: Top 73%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.