Comparative evaluation of march 2026-generation large language models for heart failure self-management education: a blinded, multidimensional in silico study
Self-management education is a cornerstone of chronic heart failure (HF) care, yet clinical workload limits its delivery. Patients increasingly use large language models (LLMs) for HF information, but comparative data on March 2026-generation models remain scarce. We compared five publicly available LLMs against expert-written HF education materials across information quality, readability, actionability, clinical accuracy, and safety. Thirty-two standardized patient questions across eight HF self-management domains were submitted in zero-shot format to five LLMs (Claude Sonnet 4.6, GPT-5.2, DeepSeek-V3.2, Kimi K2.5, Gemini 3.1 Pro). Six fully blinded raters evaluated anonymized responses using DISCERN, PEMAT-P, four readability indices (FKGL, GFI, SMOG, ARI), clinical accuracy ratings, and a three-tier error severity scale. Inter-rater reliability (ICC) and linear mixed-effects models were used to compare performance across sources, topics, and their interaction. Inter-rater reliability was good (ICC 0.77–0.84, all p < 0.001). Human expert materials outperformed all LLMs across all metrics (all adjusted p < 0.05). Among LLMs, GPT-5.2 achieved the highest estimated DISCERN scores (68.7, 95% CI 67.4–70.0), Claude Sonnet 4.6 the highest clinical accuracy (92.4%, 95% CI 91.3–93.5%), DeepSeek-V3.2 the highest PEMAT actionability (82.4, 95% CI 80.9–83.9), and Kimi K2.5 the most favorable ARI readability (7.9, 95% CI 7.5–8.3). However, pairwise differences between top-performing LLMs did not always reach statistical significance, and topic-level estimates were based on only four questions per domain. Significant source × topic interactions were observed (all p < 0.001). Total clinical errors ranged from 8 (Claude Sonnet 4.6) to 27 (Gemini 3.1 Pro); high-risk errors were distributed across all five LLMs, concentrated in warning-sign and medication domains. March 2026-generation LLMs show domain-specific promise for HF patient education, but high-risk errors persist across all systems. LLM-generated content requires mandatory clinical review, and model selection should match specific educational tasks rather than relying on generic performance rankings.
Authors
- Guanghe Li
- Jiayi Shi
- Yumei Jiang
- Juan Li
Institutions
- Tongji University (CN)
- Second Affiliated Hospital of Nanjing Medical University (CN)
- Tongji Hospital (CN)
- University of Rochester (US)
Publication Details
- Journal
- Scientific Reports
- Published
- 2026-09-25
- DOI
- https://doi.org/10.1038/s41598-026-72635-3
- Primary Topic
- Heart Failure Treatment and Management
- Type
- article
- Field-Weighted Citation Impact
- 0.00