A Large Language Model–Powered Multiagent Framework Emulating Standardized Patients in Clinical Communication Skills Training: Development and Evaluation Study
Background: Effective clinical communication is essential for medical practice, with standardized patients (SPs) being a reliable standard training method despite resource limitations. While large language models (LLMs) show strong role-playing abilities, current virtual patients (VPs) based on single LLMs face fidelity and interaction challenges. Recent advances in multiagent frameworks, which have demonstrated considerable potential in handling complex tasks, offer a new perspective for creating VPs in medical education. Objective: This study aimed to develop and evaluate a novel multiagent VP framework that simulates SPs through a collaborative agent design, thereby enhancing human-like fidelity and interaction performance in clinical communication training-oriented VP simulation. Methods: Our multiagent framework constructed 5 specialized subagents by simulating the functional partitioning of brain regions, collaboratively simulating the entire process, from case reception to interactive consultation scenarios, designed for medical students. To enhance the interaction performance of VPs, we incorporated retrieval-augmented technology, while deep character reasoning was used to improve response richness and realism. We evaluated the proposed framework through a 2-phase experiment in which the metrics of response quality, role-playing performance, interaction efficiency, information accumulation, and perceived educational utility were applied consistently: first, to compare different base models, and second, to benchmark the complete framework against a single-LLM baseline. Results: The multiagent framework outperformed single-LLM baselines across multiple evaluation settings, achieving high information accuracy and role-playing scores under standardized dialogue conditions. Specifically, the GPT-4o-based implementation achieved peak factual consistency of 0.769 (SD 0.04), while all configurations maintained >94% clinical accuracy. The Qwen3-32B-based framework achieved the lowest misleading rate of 1.28% (SD 1.20), compared to 4.72% (SD 1.53%) for single-LLM scoring. In assessments using standard dialogue scripts, the Qwen3-32B-based framework attained the highest role-playing competency score of 39.67 (SD 0.71) and received high expert praise. However, limited discriminative power against specific leading questions on low-quality inquiries indicated that while these findings specifically establish high fidelity under structured conditions, further adaptation is required for authentic student interactions. Interaction efficiency remained practical with acceptable latency (~3 s) based on Qwen3-32B while maintaining a stable information pace during multiturn dialogues. Furthermore, a preliminary exploration of factual consistency and role-playing ability across 5 clinical departments demonstrated potential scalability. Conclusions: The multiagent framework offers a viable simulation of SPs through the coordinated interaction of multiple LLM-based agents. This approach enhances the performance of VP simulation, providing a customizable and scalable solution for medical communication training, without compromising patient confidentiality. The framework holds substantial potential for advancing medical education approaches.
Authors
- Xudong Lü (ORCID: https://orcid.org/0000-0001-7658-5250)
- Yunzi Long
- Yijie Wang (ORCID: https://orcid.org/0000-0002-2023-6972)
- Yufei Qu (ORCID: https://orcid.org/0009-0004-5333-3367)
- Jiao Li (ORCID: https://orcid.org/0000-0001-6391-8343)
- Xiaowei Xu (ORCID: https://orcid.org/0000-0003-2863-1376)
Institutions
- Chinese Academy of Medical Sciences & Peking Union Medical College (CN)
- Peking University (CN)
- National Clinical Research Center for Digestive Diseases (CN)
- Hangzhou Medical College (CN)
- Zhejiang University (CN)
Publication Details
- Journal
- Journal of Medical Internet Research
- Published
- 2026-06-04
- DOI
- https://doi.org/10.2196/84747
- Primary Topic
- Artificial Intelligence in Healthcare and Education
- Type
- article
- Field-Weighted Citation Impact
- 0.00
Funders
- Zhejiang University
- Beijing Municipal Health Commission
- Chinese Academy of Medical Sciences
- Peking University
- Peking Union Medical College
- Peking University Health Science Center