RareMed-Dialogue: a multi-turn dialogue and benchmark on long-tail clinical cases
Artificial intelligence shows considerable promise for clinical decision-making. However, developing models that generalize across heterogeneous real-world settings while maintaining complex diagnostic reasoning remains challenging. There is a need for high-quality datasets that capture long-tail, multi-stage clinical reasoning processes to support robust AI deployment in clinical practice. We introduce RareMed-Dialogue, a high-quality, auditable Chinese multi-turn clinical dialogue dataset constructed from 10,114 authentic difficult, rare, and complex cases representing the clinical long tail. These cases focus on high-risk clinical scenarios characterized by structural complexity, multidisciplinary coordination, and iterative diagnostic inquiry. Each clinical case was transformed into a rule-based template covering eight key stages of the diagnostic and therapeutic workflow. Multi-turn dialogues were then generated under the guidance of multidisciplinary expert input, followed by human verification and augmentation using large language models. To evaluate dataset quality, we developed a dedicated evaluation framework incorporating two novel metrics: Dialogue-chain Integrity (DCI) and Stage Accuracy (SA), designed to assess coherence and correctness across multi-turn clinical reasoning processes. RareMed-Dialogue demonstrates three key properties: (i) multi-turn, dialogue-driven reconstruction of diagnostic and therapeutic workflows; (ii) integration of multi-stage clinical tasks within a unified conversational structure; and (iii) high clinical fidelity with explicit interpretability. The dataset faithfully reproduces complex diagnostic processes requiring iterative inquiry and multidisciplinary collaboration in long-tail clinical scenarios. RareMed-Dialogue provides a structured resource for training and benchmarking artificial intelligence systems on complex clinical reasoning tasks. Further external validation is needed to determine its generalizability to real-world clinical settings. Data access and usage details are provided in the Supplementary Materials.
Authors
- Lan Yang (ORCID: https://orcid.org/0009-0009-5494-6698)
- Qianli Zhao (ORCID: https://orcid.org/0000-0003-2844-1644)
- Yang Liu (ORCID: https://orcid.org/0000-0002-0367-8189)
- Haitao Zhang
- Dali Zhang
- Ronghao Wang
- Yiwei Liu
Institutions
- Johns Hopkins University Applied Physics Laboratory (US)
- Beijing Academy of Artificial Intelligence (CN)
- Shanghai Pulmonary Hospital (CN)
- Shanghai East Hospital (CN)
Publication Details
- Journal
- BMC Medical Informatics and Decision Making
- Published
- 2026-09-15
- DOI
- https://doi.org/10.1186/s12911-026-03846-x
- Primary Topic
- Machine Learning in Healthcare
- Type
- article
- Field-Weighted Citation Impact
- 0.00
Funders
- National Natural Science Foundation of China