Comparative evaluation of march 2026-generation large language models for heart failure self-management education: a blinded, multidimensional in silico study

Self-management education is a cornerstone of chronic heart failure (HF) care, yet clinical workload limits its delivery. Patients increasingly use large language models (LLMs) for HF information, but comparative data on March 2026-generation models remain scarce. We compared five publicly available LLMs against expert-written HF education materials across information quality, readability, actionability, clinical accuracy, and safety. Thirty-two standardized patient questions across eight HF self-management domains were submitted in zero-shot format to five LLMs (Claude Sonnet 4.6, GPT-5.2, DeepSeek-V3.2, Kimi K2.5, Gemini 3.1 Pro). Six fully blinded raters evaluated anonymized responses using DISCERN, PEMAT-P, four readability indices (FKGL, GFI, SMOG, ARI), clinical accuracy ratings, and a three-tier error severity scale. Inter-rater reliability (ICC) and linear mixed-effects models were used to compare performance across sources, topics, and their interaction. Inter-rater reliability was good (ICC 0.77–0.84, all p < 0.001). Human expert materials outperformed all LLMs across all metrics (all adjusted p < 0.05). Among LLMs, GPT-5.2 achieved the highest estimated DISCERN scores (68.7, 95% CI 67.4–70.0), Claude Sonnet 4.6 the highest clinical accuracy (92.4%, 95% CI 91.3–93.5%), DeepSeek-V3.2 the highest PEMAT actionability (82.4, 95% CI 80.9–83.9), and Kimi K2.5 the most favorable ARI readability (7.9, 95% CI 7.5–8.3). However, pairwise differences between top-performing LLMs did not always reach statistical significance, and topic-level estimates were based on only four questions per domain. Significant source × topic interactions were observed (all p < 0.001). Total clinical errors ranged from 8 (Claude Sonnet 4.6) to 27 (Gemini 3.1 Pro); high-risk errors were distributed across all five LLMs, concentrated in warning-sign and medication domains. March 2026-generation LLMs show domain-specific promise for HF patient education, but high-risk errors persist across all systems. LLM-generated content requires mandatory clinical review, and model selection should match specific educational tasks rather than relying on generic performance rankings.

Authors

Institutions

Publication Details

Journal
Scientific Reports
Published
2026-09-25
DOI
https://doi.org/10.1038/s41598-026-72635-3
Primary Topic
Heart Failure Treatment and Management
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Comparative evaluation of march 2026-generation large language models for heart failure self-management education: a blinded, multidimensional in silico study

Guanghe Li, Jiayi Shi, Yumei Jiang, Juan Li
Scientific Reports
Heart Failure Treatment and Management
article

Comparative evaluation of march 2026-generation large language models for heart failure self-management education: a blinded, multidimensional in silico study

Guanghe Li, Jiayi Shi, Yumei Jiang, Juan Li
article en

Abstract

Self-management education is a cornerstone of chronic heart failure (HF) care, yet clinical workload limits its delivery. Patients increasingly use large language models (LLMs) for HF information, but comparative data on March 2026-generation models remain scarce. We compared five publicly available LLMs against expert-written HF education materials across information quality, readability, actionability, clinical accuracy, and safety. Thirty-two standardized patient questions across eight HF self-management domains were submitted in zero-shot format to five LLMs (Claude Sonnet 4.6, GPT-5.2, DeepSeek-V3.2, Kimi K2.5, Gemini 3.1 Pro). Six fully blinded raters evaluated anonymized responses using DISCERN, PEMAT-P, four readability indices (FKGL, GFI, SMOG, ARI), clinical accuracy ratings, and a three-tier error severity scale. Inter-rater reliability (ICC) and linear mixed-effects models were used to compare performance across sources, topics, and their interaction. Inter-rater reliability was good (ICC 0.77–0.84, all p < 0.001). Human expert materials outperformed all LLMs across all metrics (all adjusted p < 0.05). Among LLMs, GPT-5.2 achieved the highest estimated DISCERN scores (68.7, 95% CI 67.4–70.0), Claude Sonnet 4.6 the highest clinical accuracy (92.4%, 95% CI 91.3–93.5%), DeepSeek-V3.2 the highest PEMAT actionability (82.4, 95% CI 80.9–83.9), and Kimi K2.5 the most favorable ARI readability (7.9, 95% CI 7.5–8.3). However, pairwise differences between top-performing LLMs did not always reach statistical significance, and topic-level estimates were based on only four questions per domain. Significant source × topic interactions were observed (all p < 0.001). Total clinical errors ranged from 8 (Claude Sonnet 4.6) to 27 (Gemini 3.1 Pro); high-risk errors were distributed across all five LLMs, concentrated in warning-sign and medication domains. March 2026-generation LLMs show domain-specific promise for HF patient education, but high-risk errors persist across all systems. LLM-generated content requires mandatory clinical review, and model selection should match specific educational tasks rather than relying on generic performance rankings.

Scientific Reports
Tongji University (CN), Second Affiliated Hospital of Nanjing Medical University (CN), Tongji Hospital (CN), University of Rochester (US)
Quality Education
Openalex Percentile: Top 11%
Heart Failure Treatment and Management
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.