Automated Health Care Thematic Analysis Using a Multiagent Large Language Model: Algorithm Development and Evaluation Study
Background Understanding patients’ experiences is essential for advancing patient-centered care, especially in chronic diseases that require ongoing communication. Qualitative thematic analysis is widely used to explore these experiences; however, the process remains labor-intensive, subjective, and difficult to scale. Objective This study aimed to develop and evaluate Collaborative Theme Identification Agents (CoTI), a multiagent large language model framework designed to support manual thematic analysis by rapidly generating supporting excerpts, initial codes, and themes. Methods CoTI consists of 3 agents: Instructor, Thematizer, and CodebookGenerator. The Instructor refines instruction prompts, the Thematizer extracts supporting excerpts and generates initial codes for each transcript, and the CodebookGenerator groups similar codes across all transcripts into a codebook with themes. We evaluated CoTI primarily using 12 transcripts of patient with heart failure, with a focus on perceptions of medication intensity. CoTI-generated outputs were compared against the reference standard developed by senior investigators. To explore human-AI interaction in thematic analyses, we further implemented CoTI in a user-facing application. Results CoTI generated supporting excerpts, initial codes, and themes that were more similar to those of senior investigators than were the outputs of junior investigators, baseline natural language processing models, and other basic large language models. In an exploratory human-AI collaboration experiment, we found that the collaboration between CoTI and junior investigators provided only marginal gains compared to CoTI alone. A possible hypothesis was that junior investigators may overrely on CoTI and limit their independent critical thinking. Conclusions CoTI can improve the efficiency of thematic analysis by rapidly generating supporting excerpts, initial codes, and themes for human researchers’ review. These findings highlight CoTI’s potential as a useful tool for scalable qualitative research.
Authors
- Alexander Wen (ORCID: https://orcid.org/0009-0008-9038-3039)
- Min Ji Kwak (ORCID: https://orcid.org/0000-0003-2778-3984)
- De’Angelo Hermesky
- Yejin Kim (ORCID: https://orcid.org/0000-0001-7815-6310)
- Qidi Xu (ORCID: https://orcid.org/0009-0006-6230-2719)
- Alexa Cumming
- Nuzha Amjad (ORCID: https://orcid.org/0009-0005-1821-2807)
- Grace Giles (ORCID: https://orcid.org/0009-0008-0824-0128)
Publication Details
- Journal
- Journal of Medical Internet Research
- Published
- 2026-09-30
- DOI
- https://doi.org/10.2196/90872
- Primary Topic
- Machine Learning in Healthcare
- Type
- article
- Field-Weighted Citation Impact
- 0.00