On-premise medical AI agents for reliable clinical decision-making

Abstract Autonomous clinical artificial intelligence (AI) agents powered by large language models (LLMs), meaning systems that can complete a diagnostic workflow without continuous human input, are increasingly capable of supporting complex reasoning and decision-making. Clinical translation, however, remains limited by two unmet requirements: institutionally governed deployment and reliable decision-time uncertainty estimation. Here we developed and evaluated a fully on-premise clinical agent that couples local operational control with a multi-perspective reliability framework to support selective autonomy. Across two Medical Information Mart for Intensive Care IV (MIMIC-IV)-derived benchmarks, the agent achieved 90.04% accuracy on a seven-disease task and 83.8% accuracy on a four-disease task, approaching a cloud baseline on the primary benchmark. To assess decision-time reliability, we quantified internal-likelihood, language-based and behavioral-stability measures for diagnosis and reasoning. Diagnostic behavioral consistency provided the strongest discrimination of correctness (area under the curve (AUC) = 0.860) and remained informative under stress testing (AUC = 0.875). At a consistency threshold of 0.90, 49.4% of cases were retained at 98.9% diagnostic accuracy. These findings support a practical framework for institutionally governed clinical agents in which decision-time reliability signals identify a lower-risk subset for autonomous handling and defer the remainder for review.

Authors

Institutions

Publication Details

Journal
Nature Medicine
Published
2026-09-15
DOI
https://doi.org/10.1038/s41591-026-04609-x
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

On-premise medical AI agents for reliable clinical decision-making

Fabian Wolf, Jan Clusmann, Lino Möhrmann, Dyke Ferber et al.
Nature Medicine
Artificial Intelligence in Healthcare and Education
article

On-premise medical AI agents for reliable clinical decision-making

Fabian Wolf, Jan Clusmann, Lino Möhrmann, Dyke Ferber, Georg Wölflein, Catharina Wichmann, Jakob Nikolas Kather, Xuewei Wu, Elena E. Möhrmann, Li Zhang, Zunamys I. Carrero, Julien Vibert, Junhao Liang, Tim Lenz
article en

Abstract

Abstract Autonomous clinical artificial intelligence (AI) agents powered by large language models (LLMs), meaning systems that can complete a diagnostic workflow without continuous human input, are increasingly capable of supporting complex reasoning and decision-making. Clinical translation, however, remains limited by two unmet requirements: institutionally governed deployment and reliable decision-time uncertainty estimation. Here we developed and evaluated a fully on-premise clinical agent that couples local operational control with a multi-perspective reliability framework to support selective autonomy. Across two Medical Information Mart for Intensive Care IV (MIMIC-IV)-derived benchmarks, the agent achieved 90.04% accuracy on a seven-disease task and 83.8% accuracy on a four-disease task, approaching a cloud baseline on the primary benchmark. To assess decision-time reliability, we quantified internal-likelihood, language-based and behavioral-stability measures for diagnosis and reasoning. Diagnostic behavioral consistency provided the strongest discrimination of correctness (area under the curve (AUC) = 0.860) and remained informative under stress testing (AUC = 0.875). At a consistency threshold of 0.90, 49.4% of cases were retained at 98.9% diagnostic accuracy. These findings support a practical framework for institutionally governed clinical agents in which decision-time reliability signals identify a lower-risk subset for autonomous handling and defer the remainder for review.

Nature Medicine
University of Leeds (GB), Inserm (FR), Heidelberg University (DE), Université Paris-Saclay (FR), Institut Gustave Roussy (FR), University Hospital Heidelberg (DE), National Center for Tumor Diseases (DE), First Affiliated Hospital of Jinan University (CN), Hochschule für Technik und Wirtschaft Dresden – University of Applied Sciences (DE), Else Kröner Fresenius Center for Digital Health (DE), Technische Universität Dresden (DE)
Peace, Justice and strong institutions
Openalex Percentile: Top 14%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.