The complexity paradox of generalist AI agents in medicine: exploring GPT-5 for real-world multimodal medical diagnosis

The next generation of medical AI systems that integrate language understanding, web data retrieval, and multi-step reasoning are demonstrating increasing promise as potential diagnostic aids in technical evaluations. However, their real world diagnostic performance, explainability, and reliability remain unclear. We conducted a systematic quantitative analysis of ten GPT-5-based configurations on 161 clinician validated real world cases. Each model was assessed for diagnostic accuracy, reasoning quality, retrieval accuracy, evidence integration, and robustness to low-quality inputs. We found important limitations in these AI systems, including image misinterpretation, inconsistent web search that sometimes degraded performance, and limited transparency in agentic workflows. Yet the models also showed strengths, including effective use of patient history, clear stepwise reasoning, plausible differentials, improved specificity with explicit reasoning, and relevant evidence retrieval. To support clinical deployment and future model development, we recommend prioritizing reasoning-optimized AI models, improving information retrieval mechanisms for reflection and error correction, strengthening visual-textual integration, and establishing practical standards for image quality and clarity. Overall, reliable real world multimodal diagnostic performance depends more on transparent reasoning and high-quality inputs than on aggressive retrieval or complex agentic workflows. Persistent gaps in evidence integration and visual reasoning remain major barriers to safe routine clinical use.

Authors

Institutions

Publication Details

Journal
npj Digital Medicine
Published
2026-10-09
DOI
https://doi.org/10.1038/s41746-026-03325-7
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

The complexity paradox of generalist AI agents in medicine: exploring GPT-5 for real-world multimodal medical diagnosis

Jing Yang Huang, Yonghui Wu, Adam Kiridly, Lifang He et al.
npj Digital Medicine
Artificial Intelligence in Healthcare and Education
article

The complexity paradox of generalist AI agents in medicine: exploring GPT-5 for real-world multimodal medical diagnosis

Jing Yang Huang, Yonghui Wu, Adam Kiridly, Lifang He, Omar Toubat, Gregory E. Tasian, Michael A. Catalano, Wei Liu, Lei Xing, Lichao Sun, Jiarong Qian, Zhiling Yan, Hua Xu, Shaohui Zhang, Quanzheng Li, Zhiyong Lu, Kai Zhang, Xiang Li
article en

Abstract

The next generation of medical AI systems that integrate language understanding, web data retrieval, and multi-step reasoning are demonstrating increasing promise as potential diagnostic aids in technical evaluations. However, their real world diagnostic performance, explainability, and reliability remain unclear. We conducted a systematic quantitative analysis of ten GPT-5-based configurations on 161 clinician validated real world cases. Each model was assessed for diagnostic accuracy, reasoning quality, retrieval accuracy, evidence integration, and robustness to low-quality inputs. We found important limitations in these AI systems, including image misinterpretation, inconsistent web search that sometimes degraded performance, and limited transparency in agentic workflows. Yet the models also showed strengths, including effective use of patient history, clear stepwise reasoning, plausible differentials, improved specificity with explicit reasoning, and relevant evidence retrieval. To support clinical deployment and future model development, we recommend prioritizing reasoning-optimized AI models, improving information retrieval mechanisms for reflection and error correction, strengthening visual-textual integration, and establishing practical standards for image quality and clarity. Overall, reliable real world multimodal diagnostic performance depends more on transparent reasoning and high-quality inputs than on aggressive retrieval or complex agentic workflows. Persistent gaps in evidence integration and visual reasoning remain major barriers to safe routine clinical use.

npj Digital Medicine
National Institutes of Health (US), Mayo Clinic (US), Children's Hospital of Philadelphia (US), Lehigh University (US), United States National Library of Medicine (US), University of Florida Health (US), Yale University (US), University of Florida (US), Massachusetts General Hospital (US), Stanford Medicine (US), Rice University (US), University of Pennsylvania (US), Stanford University (US)
Openalex Percentile: Top 19%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.