OsteoCHAT: real-world patient evaluation and benchmarking of a guideline-grounded osteoporosis chatbot

Patients with osteoporosis have information needs that routine care inconsistently addresses. In this multicenter real-world evaluation, the LLM-based, guideline-grounded chatbot OsteoCHAT showed high participant acceptability and achieved expert-rated answer quality comparable to general-purpose LLMs. Despite challenges in complex treatment and safety-related topics, guideline-grounded chatbots may represent useful educational adjuncts. Patients with osteoporosis have information needs that routine care inconsistently addresses. Guideline-grounded chatbots may offer scalable educational support We developed OsteoCHAT, a retrieval-augmented conversational system based on the current German osteoporosis guideline, and conducted a multicenter evaluation at seven centers in Germany. Participants interacted with OsteoCHAT, provided real-time feedback on individual responses, and completed a standardized questionnaire assessing usability, comprehensibility, usefulness, trust, perceived quality, and preference over conventional internet search. In parallel, OsteoCHAT was benchmarked against three general-purpose large language models (LLMs) using ten osteoporosis frequently asked questions derived from the BfO patient guideline as the gold standard. Five clinical experts independently rated these responses across five domains, each scored 0–3. Overall, 1,075 question-answer interactions were recorded. Of 389 individually rated responses, 379 (97.4%) received positive feedback. Among 217 complete questionnaire responses, 84.3% rated answers as easily understandable, 79.8% found OsteoCHAT easy to use, 77.2% perceived time savings, and 78.7% considered it a useful addition to existing educational materials. Trust was comparatively lower (64.7% agreement). Blinded expert benchmarking showed acceptable-to-high-quality responses across all LLMs (median total scores 14/15 for Gemini, ChatGPT, and OsteoCHAT and 13/15 for Meta AI), with expert-identified weaknesses mainly concerning pharmacological treatment indication and communication of adverse and rare safety-critical events. OsteoCHAT demonstrated high user acceptability and usability in a real-world multicenter evaluation, was rated as a valuable addition to existing educational materials, and achieved expert-rated response quality comparable to general-purpose LLMs, although limitations in safety communication were identified across all examined systems.

Authors

Institutions

Publication Details

Journal
Osteoporosis International
Published
2026-10-03
DOI
https://doi.org/10.1007/s00198-026-08247-4
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

OsteoCHAT: real-world patient evaluation and benchmarking of a guideline-grounded osteoporosis chatbot

Karin Mahn, Björn Bühring, Asarnusch Rashid, Martin Gehlen et al.
Osteoporosis International
Artificial Intelligence in Healthcare and Education
article

OsteoCHAT: real-world patient evaluation and benchmarking of a guideline-grounded osteoporosis chatbot

Karin Mahn, Björn Bühring, Asarnusch Rashid, Martin Gehlen, Paula Hoff, Frank Buttgereit, Edgar Wiebe, Olga Seifert, Nils Schulz, Peter Oelzner, Uwe Lange, Alexander Pfeil, Sebastian Kuhn, Philipp Klemm, Johannes Knitza
article en

Abstract

Patients with osteoporosis have information needs that routine care inconsistently addresses. In this multicenter real-world evaluation, the LLM-based, guideline-grounded chatbot OsteoCHAT showed high participant acceptability and achieved expert-rated answer quality comparable to general-purpose LLMs. Despite challenges in complex treatment and safety-related topics, guideline-grounded chatbots may represent useful educational adjuncts. Patients with osteoporosis have information needs that routine care inconsistently addresses. Guideline-grounded chatbots may offer scalable educational support We developed OsteoCHAT, a retrieval-augmented conversational system based on the current German osteoporosis guideline, and conducted a multicenter evaluation at seven centers in Germany. Participants interacted with OsteoCHAT, provided real-time feedback on individual responses, and completed a standardized questionnaire assessing usability, comprehensibility, usefulness, trust, perceived quality, and preference over conventional internet search. In parallel, OsteoCHAT was benchmarked against three general-purpose large language models (LLMs) using ten osteoporosis frequently asked questions derived from the BfO patient guideline as the gold standard. Five clinical experts independently rated these responses across five domains, each scored 0–3. Overall, 1,075 question-answer interactions were recorded. Of 389 individually rated responses, 379 (97.4%) received positive feedback. Among 217 complete questionnaire responses, 84.3% rated answers as easily understandable, 79.8% found OsteoCHAT easy to use, 77.2% perceived time savings, and 78.7% considered it a useful addition to existing educational materials. Trust was comparatively lower (64.7% agreement). Blinded expert benchmarking showed acceptable-to-high-quality responses across all LLMs (median total scores 14/15 for Gemini, ChatGPT, and OsteoCHAT and 13/15 for Meta AI), with expert-identified weaknesses mainly concerning pharmacological treatment indication and communication of adverse and rare safety-critical events. OsteoCHAT demonstrated high user acceptability and usability in a real-world multicenter evaluation, was rated as a valuable addition to existing educational materials, and achieved expert-rated response quality comparable to general-purpose LLMs, although limitations in safety communication were identified across all examined systems.

Osteoporosis International
Philipps University of Marburg (DE), Justus-Liebig-Universität Gießen (DE), Laboklin (Germany) (DE), German Rheumatism Research Centre (DE), University Hospital Leipzig (DE), St. Josef Krankenhaus (DE), Jena University Hospital (DE), Klinik und Poliklinik für Neurologie (DE), Endokrinologikum (DE), Stiftung Charité, Friedrich Schiller University Jena (DE), Charité - Universitätsmedizin Berlin (DE)
Openalex Percentile: Top 15%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.