Development and Nationwide Multicentre Evaluation of Guideline-Grounded Large Language Model Chatbots to Support Patient Self-Management and Education in Rheumatology

Patients with rheumatic diseases have persistent information needs that are not fully addressed in routine care. We developed and evaluated guideline-grounded, large language model (LLM) chatbots to support patient self-management and education in rheumatology.Ten disease-specific chatbots based on German guidelines were co-developed and deployed through 13 rheumatology centres and six patient organisations. Chatbot users rated responses and completed a questionnaire. User questions, feedback, and response characteristics were analysed using category-based coding and a six-dimensional LLM-as-a-judge assessment, with LLM-based ratings compared with rheumatologist ratings in random subsets.Between September 2025 and January 2026, 6291 questions were recorded. Thirteen question categories were identified, most commonly disease-specific questions (50.2%), medication and monitoring (37.3%) and diagnostics (28.2%). The chatbots were unable to answer in 263 interactions (4.2%). Of 2671 responses rated by users, 2481 (92.9%) received a positive rating. Insufficient detail was the most common reason for negative ratings (125/190, 65.8%). Among 602 questionnaire respondents, 84.6% reported that the chatbot was easy to use, 84.1% that answers were easy to understand, and 80.2% that it was a useful addition to patient education. In the LLM-based evaluation, 95.3% of answers were rated as completely safe and 79.1% as completely correct. Guideline adherence was assessed separately, with 45.0% rated as fully adherent; agreement with physician assessment was weak.Guideline-grounded chatbots received predominantly positive user feedback in real-world use, while LLM-based evaluation suggested that most responses were safe and correct. User questions and feedback may help guide iterative improvements to source content and patient education materials. Further studies are needed to evaluate educational effectiveness and independently validate response quality and clinical safety.

Authors

Institutions

Publication Details

Journal
Journal of Medical Systems
Published
2026-10-05
DOI
https://doi.org/10.1007/s10916-026-02470-6
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Development and Nationwide Multicentre Evaluation of Guideline-Grounded Large Language Model Chatbots to Support Patient Self-Management and Education in Rheumatology

Harriet Morf, Ulrich Drott, MD Peer Aries, Philipp Klemm et al.
Journal of Medical Systems
Artificial Intelligence in Healthcare and Education
article

Development and Nationwide Multicentre Evaluation of Guideline-Grounded Large Language Model Chatbots to Support Patient Self-Management and Education in Rheumatology

Harriet Morf, Ulrich Drott, MD Peer Aries, Philipp Klemm, Asarnusch Rashid, Axel J. Hueber, Johannes Hornig, Diego Benavent, Vanessa Bartsch, Johannes Knitza, Alexander Pfeil, Peter Böhm, Felix Mühlensiepen, Tim Wilhelmi, Sebastian Kuhn, Martin Krusche, Hannah Labinsky, Marius Platt, Marie-Therese Holzer, Max Müller, Gabriel Dischereit, Daniel Fink
article en

Abstract

Patients with rheumatic diseases have persistent information needs that are not fully addressed in routine care. We developed and evaluated guideline-grounded, large language model (LLM) chatbots to support patient self-management and education in rheumatology.Ten disease-specific chatbots based on German guidelines were co-developed and deployed through 13 rheumatology centres and six patient organisations. Chatbot users rated responses and completed a questionnaire. User questions, feedback, and response characteristics were analysed using category-based coding and a six-dimensional LLM-as-a-judge assessment, with LLM-based ratings compared with rheumatologist ratings in random subsets.Between September 2025 and January 2026, 6291 questions were recorded. Thirteen question categories were identified, most commonly disease-specific questions (50.2%), medication and monitoring (37.3%) and diagnostics (28.2%). The chatbots were unable to answer in 263 interactions (4.2%). Of 2671 responses rated by users, 2481 (92.9%) received a positive rating. Insufficient detail was the most common reason for negative ratings (125/190, 65.8%). Among 602 questionnaire respondents, 84.6% reported that the chatbot was easy to use, 84.1% that answers were easy to understand, and 80.2% that it was a useful addition to patient education. In the LLM-based evaluation, 95.3% of answers were rated as completely safe and 79.1% as completely correct. Guideline adherence was assessed separately, with 45.0% rated as fully adherent; agreement with physician assessment was weak.Guideline-grounded chatbots received predominantly positive user feedback in real-world use, while LLM-based evaluation suggested that most responses were safe and correct. User questions and feedback may help guide iterative improvements to source content and patient education materials. Further studies are needed to evaluate educational effectiveness and independently validate response quality and clinical safety.

Journal of Medical SystemsVol. 50(1)
Philipps University of Marburg (DE), Friedrich-Alexander-Universität Erlangen-Nürnberg (DE), Justus-Liebig-Universität Gießen (DE), University of Würzburg (DE), Bellvitge University Hospital (ES), Humboldt-Universität zu Berlin (DE), Laboklin (Germany) (DE), Universitätsklinikum Erlangen (DE), Nuremberg Hospital (DE), University Medical Center Hamburg-Eppendorf (DE), Deutsche Rheuma-Liga (DE), Universitätsklinikum Würzburg (DE), Jena University Hospital (DE), Friedrich Schiller University Jena (DE), Charité - Universitätsmedizin Berlin (DE)
Deutsche Rheumastiftung, HORIZON EUROPE Framework Programme
Good health and well-being
Openalex Percentile: Top 19%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.