Early evidence for a multi-agent AI simulator for clinical reasoning practice: performance, consistency, and challenges

OBJECTIVES: Clinical reasoning develops through repeated, deliberate practice, yet clinical and simulation environments are often limited by continuity, feedback, and scalability constraints. Large language models (LLMs) may address this by generating virtual clinical encounters accessible without live instructors or standardized patients but validity evidence remains limited. This study explored the performance of MAESSCR (Multi-Agent Educational Scenario Simulator for Clinical Reasoning), a platform where multiple LLM-based agents, each assigned a distinct role (i.e. patient, physical exam, diagnostic testing) interact with learners through a text-based encounter. METHODS: In Fall of 2024, 175 second-year medical students completed three MAESSCR clinical encounters as coursework. Six clinician-educators developed a rating tool to evaluate 120 transcripts (40 randomly sampled per case) using a dichotomous (yes/no) scale across four domains: (1) realism of agent responses, (2) adherence to scripted case details, (3) platform functionality, and (4) interference with students' independent reasoning through clinical findings. Qualitative narrative review supplemented binary ratings. RESULTS: MAESSCR followed scripted details in 92 % (110/120) of transcripts. Diagnosis-changing information occurred in 2.5 % (3/120). Unrealistic patient portrayal appeared in 12 % (14/120) of encounters, and technical issues in 17 % (20/120). AI agents most often interfered with student's clinical reasoning by interpreting findings before the students had the opportunity. This occurred in 28 % (34/120) of encounters due to history/physical exam agents and 59 % (71/120) because of diagnostics/management agents. Most disruptions were minor and unlikely to compromise the overall encounter. CONCLUSIONS: Multi-agent, LLM-based simulations offer a scalable approach to deliberate practice of clinical reasoning. Educational value depends on role stability, contextual fidelity, and learners' opportunity to interpret clinical information independently. Specific design considerations are essential to ensure AI-generated simulations support, rather than disrupt, the clinical reasoning process.

Authors

Institutions

Publication Details

Journal
Diagnosis
Published
2026-09-15
DOI
https://doi.org/10.1515/dx-2026-0069
Primary Topic
Simulation-Based Education in Healthcare
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Early evidence for a multi-agent AI simulator for clinical reasoning practice: performance, consistency, and challenges

Weibing Zheng, James Bowen, Laurah Turner, Matthew Kelleher et al.
Diagnosis
Simulation-Based Education in Healthcare
article

Early evidence for a multi-agent AI simulator for clinical reasoning practice: performance, consistency, and challenges

Weibing Zheng, James Bowen, Laurah Turner, Matthew Kelleher, Sally A. Santen, Danielle E. Weber, Seth W Overla, Jose Generoso, Christine Y. Zhou
article en

Abstract

OBJECTIVES: Clinical reasoning develops through repeated, deliberate practice, yet clinical and simulation environments are often limited by continuity, feedback, and scalability constraints. Large language models (LLMs) may address this by generating virtual clinical encounters accessible without live instructors or standardized patients but validity evidence remains limited. This study explored the performance of MAESSCR (Multi-Agent Educational Scenario Simulator for Clinical Reasoning), a platform where multiple LLM-based agents, each assigned a distinct role (i.e. patient, physical exam, diagnostic testing) interact with learners through a text-based encounter. METHODS: In Fall of 2024, 175 second-year medical students completed three MAESSCR clinical encounters as coursework. Six clinician-educators developed a rating tool to evaluate 120 transcripts (40 randomly sampled per case) using a dichotomous (yes/no) scale across four domains: (1) realism of agent responses, (2) adherence to scripted case details, (3) platform functionality, and (4) interference with students' independent reasoning through clinical findings. Qualitative narrative review supplemented binary ratings. RESULTS: MAESSCR followed scripted details in 92 % (110/120) of transcripts. Diagnosis-changing information occurred in 2.5 % (3/120). Unrealistic patient portrayal appeared in 12 % (14/120) of encounters, and technical issues in 17 % (20/120). AI agents most often interfered with student's clinical reasoning by interpreting findings before the students had the opportunity. This occurred in 28 % (34/120) of encounters due to history/physical exam agents and 59 % (71/120) because of diagnostics/management agents. Most disruptions were minor and unlikely to compromise the overall encounter. CONCLUSIONS: Multi-agent, LLM-based simulations offer a scalable approach to deliberate practice of clinical reasoning. Educational value depends on role stability, contextual fidelity, and learners' opportunity to interpret clinical information independently. Specific design considerations are essential to ensure AI-generated simulations support, rather than disrupt, the clinical reasoning process.

Diagnosis
Cincinnati Children's Hospital Medical Center (US), University of Cincinnati (US), University of Cincinnati Medical Center (US)
Quality Education
Openalex Percentile: Top 11%
Simulation-Based Education in Healthcare
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.