Multidimensional evaluation of generative AI chatbots for public consultation on sudden cardiac death: safety, accuracy, empathy, reliability, quality, and readability across six models

Generative AI chatbots are increasingly used for health information, but their performance in safety-sensitive sudden cardiac death (SCD) consultation remains insufficiently characterized. This study evaluated six publicly accessible chatbots on SCD-related public questions involving symptom triage, emergency response, CPR/AED use, inherited risk, myocarditis-related concerns, return to exercise, screening, and ICD decision-making. This cross-sectional, comparative, text-level study evaluated ChatGPT Plus 5.5 thinking, Gemini 3.5 Flash, Microsoft Copilot smart, DeepSeek-V4-pro, Doubao, and Perplexity. A standardized set of 56 SCD-related public consultation questions was developed from Google Trends, online forums, public question-answer platforms, clinical guidelines, expert consensus statements, systematic reviews, and qualitative interviews. Each question was submitted once to each chatbot between May 10 and May 16, 2026, yielding 336 model-response items. Five blinded senior cardiology raters assessed safety, accuracy, empathy, DISCERN, EQIP, JAMA benchmark criteria, and GQS. Readability was calculated using six formula-based indices. Between-model comparisons were performed using Cochran’s Q test and Friedman tests, with Kendall’s W reported as the effect size. Responses containing potential safety concerns were additionally examined qualitatively. Inter-rater agreement was high across all manually assessed metrics, with Fleiss’ kappa of 0.856 for Safety and ICC values ranging from 0.801 to 0.887 for other domains. Safe responses predominated across all models, although responses containing potential safety concerns occurred in every chatbot. ChatGPT Plus 5.5 thinking and Perplexity had the highest observed safety proportions during this query window, whereas Doubao had the lowest. Qualitative analysis identified recurrent concerns involving delayed emergency activation for chest pain, syncope, or post-viral symptoms; over-reassurance about prevention; fixed return-to-exercise timelines after infection or COVID-19; pulse-check instructions that could delay CPR; and oversimplified screening, electrolyte, or ICD-related advice. Most flagged responses contained a single potential safety concern. The qualitative severity distribution was concentrated in low and low-to-moderate concerns, whereas higher-severity concerns were less frequent. Accuracy differed significantly across models ( P < 0.001; Kendall’s W = 0.864), with ChatGPT Plus 5.5 thinking, Gemini 3.5 Flash, Microsoft Copilot smart, DeepSeek-V4-pro, and Perplexity achieving similar median scores and Doubao performing lower. Empathy also differed significantly ( P < 0.001; Kendall’s W = 0.531), with DeepSeek-V4-pro having the highest observed score. Significant differences were observed for DISCERN, EQIP, JAMA, GQS, and all readability indices (all P < 0.001). Within this dataset, DeepSeek-V4-pro had the highest DISCERN and EQIP scores and the most favorable readability profile, Gemini 3.5 Flash showed comparable EQIP performance, and Microsoft Copilot smart had the highest JAMA score. However, transparency remained limited overall, and all models produced responses with relatively high reading-grade requirements. In this multidimensional evaluation of six generative AI chatbots for SCD-related public consultation, most responses were safe and clinically relevant, but potential safety concerns occurred across all models. Important limitations involved emergency escalation, CPR/AED guidance, return-to-exercise advice, individualized risk interpretation, transparency, and readability. These tools may support general SCD education, but they should not replace emergency medical services or clinician assessment. Because each question was sampled once during a defined week and all prompts were submitted in English through public interfaces accessed in China, the findings should be interpreted as a time-, language-, and access-specific snapshot rather than a stable ranking of the underlying models. Future development should prioritize evidence-linked content, clear red-flag escalation, plain-language communication, and ongoing expert review.

Authors

Institutions

Publication Details

Journal
BMC Public Health
Published
2026-10-06
DOI
https://doi.org/10.1186/s12889-026-29766-z
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Multidimensional evaluation of generative AI chatbots for public consultation on sudden cardiac death: safety, accuracy, empathy, reliability, quality, and readability across six models

Duoxue Chen, Rongyan Jiang, Youyou Chen, Haiyan Wang et al.
BMC Public Health
Artificial Intelligence in Healthcare and Education
article

Multidimensional evaluation of generative AI chatbots for public consultation on sudden cardiac death: safety, accuracy, empathy, reliability, quality, and readability across six models

Duoxue Chen, Rongyan Jiang, Youyou Chen, Haiyan Wang, Hui Ma, Huimin Wang
article en

Abstract

Generative AI chatbots are increasingly used for health information, but their performance in safety-sensitive sudden cardiac death (SCD) consultation remains insufficiently characterized. This study evaluated six publicly accessible chatbots on SCD-related public questions involving symptom triage, emergency response, CPR/AED use, inherited risk, myocarditis-related concerns, return to exercise, screening, and ICD decision-making. This cross-sectional, comparative, text-level study evaluated ChatGPT Plus 5.5 thinking, Gemini 3.5 Flash, Microsoft Copilot smart, DeepSeek-V4-pro, Doubao, and Perplexity. A standardized set of 56 SCD-related public consultation questions was developed from Google Trends, online forums, public question-answer platforms, clinical guidelines, expert consensus statements, systematic reviews, and qualitative interviews. Each question was submitted once to each chatbot between May 10 and May 16, 2026, yielding 336 model-response items. Five blinded senior cardiology raters assessed safety, accuracy, empathy, DISCERN, EQIP, JAMA benchmark criteria, and GQS. Readability was calculated using six formula-based indices. Between-model comparisons were performed using Cochran’s Q test and Friedman tests, with Kendall’s W reported as the effect size. Responses containing potential safety concerns were additionally examined qualitatively. Inter-rater agreement was high across all manually assessed metrics, with Fleiss’ kappa of 0.856 for Safety and ICC values ranging from 0.801 to 0.887 for other domains. Safe responses predominated across all models, although responses containing potential safety concerns occurred in every chatbot. ChatGPT Plus 5.5 thinking and Perplexity had the highest observed safety proportions during this query window, whereas Doubao had the lowest. Qualitative analysis identified recurrent concerns involving delayed emergency activation for chest pain, syncope, or post-viral symptoms; over-reassurance about prevention; fixed return-to-exercise timelines after infection or COVID-19; pulse-check instructions that could delay CPR; and oversimplified screening, electrolyte, or ICD-related advice. Most flagged responses contained a single potential safety concern. The qualitative severity distribution was concentrated in low and low-to-moderate concerns, whereas higher-severity concerns were less frequent. Accuracy differed significantly across models ( P < 0.001; Kendall’s W = 0.864), with ChatGPT Plus 5.5 thinking, Gemini 3.5 Flash, Microsoft Copilot smart, DeepSeek-V4-pro, and Perplexity achieving similar median scores and Doubao performing lower. Empathy also differed significantly ( P < 0.001; Kendall’s W = 0.531), with DeepSeek-V4-pro having the highest observed score. Significant differences were observed for DISCERN, EQIP, JAMA, GQS, and all readability indices (all P < 0.001). Within this dataset, DeepSeek-V4-pro had the highest DISCERN and EQIP scores and the most favorable readability profile, Gemini 3.5 Flash showed comparable EQIP performance, and Microsoft Copilot smart had the highest JAMA score. However, transparency remained limited overall, and all models produced responses with relatively high reading-grade requirements. In this multidimensional evaluation of six generative AI chatbots for SCD-related public consultation, most responses were safe and clinically relevant, but potential safety concerns occurred across all models. Important limitations involved emergency escalation, CPR/AED guidance, return-to-exercise advice, individualized risk interpretation, transparency, and readability. These tools may support general SCD education, but they should not replace emergency medical services or clinician assessment. Because each question was sampled once during a defined week and all prompts were submitted in English through public interfaces accessed in China, the findings should be interpreted as a time-, language-, and access-specific snapshot rather than a stable ranking of the underlying models. Future development should prioritize evidence-linked content, clear red-flag escalation, plain-language communication, and ongoing expert review.

BMC Public Health
University of Science and Technology of China (CN), Bozhou People's Hospital (CN)
Openalex Percentile: Top 19%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.