Assisting patients with reliable information gathering: A guide to building efficient QA systems for medical use
The application of question-answering (QA) systems in the medical domain has rapidly advanced, significantly improving patients' access to reliable health-related information. However, current approaches face notable challenges, including the difficulty in obtaining large-scale and unbiased medical datasets, significant privacy concerns, and inefficiencies due to manual dataset annotation. To address these issues, This study introduce a novel methodology leveraging publicly accessible online health forums to systematically build an unbiased, privacy-conscious QA dataset, and it employs Topic-guided Semantic Modeling (TGSM) for automated topic identification, enabling efficient and targeted annotation of relevant patient-generated content. Subsequently, this study propose a two-stage QA pipeline based on a Retriever-Reader architecture, which is further enhanced through fine-tuning state-of-the-art transformer-based models on the constructed domain-specific QA dataset. Experimental results demonstrate that our fine-tuned BioBERT significantly outperforms existing benchmarks, offering accurate patient-derived insights and providing a replicable framework for building efficient medical QA systems.
Authors
- Zichong Wang (ORCID: https://orcid.org/0000-0001-6091-6609)
- Jun Liu (ORCID: https://orcid.org/0000-0003-3808-4599)
- Wenbin Zhang (ORCID: https://orcid.org/0000-0002-4927-9570)
- Xin Ning
- Zhipeng Yin
- Min Chen
- Ian Stockwell
Institutions
- Florida International University (US)
- Carnegie Mellon University (US)
- University of Maryland, Baltimore County (US)
Publication Details
- Journal
- PLOS Digital Health
- Published
- 2026-09-17
- DOI
- https://doi.org/10.1371/journal.pdig.0001634
- Primary Topic
- Topic Modeling
- Type
- article
- Field-Weighted Citation Impact
- 0.00