Quality and Readability of AI-Generated Post-Myocardial Infarction Patient Information

Objective: Large language models are increasingly used to obtain health-related information; however, the quality and readability of their responses to practical questions after myocardial infarction (MI) remain uncertain. This study compared ChatGPT and Google Gemini with guideline-based responses for post-MI patient information.Methods: Ten commonly searched questions regarding post-MI recovery and secondary prevention were evaluated. Identical questions were submitted to ChatGPT and Google Gemini, and corresponding reference responses were derived from contemporary European Society of Cardiology guidelines. Two blinded cardiologists independently evaluated response quality using the CLEAR Tool. Readability was assessed using the Flesch Reading Ease Score (FRES) and Flesch–Kincaid Grade Level (FKGL).Results: Thirty responses were evaluated. Inter-rater reliability was excellent (ICC=0.916; 95% CI, 0.821–0.960). Total CLEAR scores differed significantly among the three sources (P=0.002). The median total CLEAR score was highest for ChatGPT [24.00 (23.63–24.38)], followed by guideline-based responses [23.25 (23.00–23.88)] and Google Gemini [19.00 (18.00–20.13)]. ChatGPT and guideline-based responses did not differ significantly in total CLEAR score (P=0.164), whereas Google Gemini scored significantly lower than both. Readability also differed significantly among sources (both FRES and FKGL, P<0.001). Google Gemini had the highest FRES [48.8 (41.0–51.6)] and lowest FKGL [11.5 (10.8–12.1)], indicating greater readability.Conclusion: ChatGPT provided post-MI information with overall quality comparable to guideline-based responses, whereas Google Gemini generated more readable but lower-quality responses. These findings suggest a potential trade-off between readability and informational quality and support the use of LLMs as adjuncts rather than substitutes for clinician-led post-MI patient education.

Authors

Institutions

Publication Details

Journal
DAHUDER Medical Journal
Published
2026-09-25
DOI
https://doi.org/10.56016/dahudermj.2015822
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Quality and Readability of AI-Generated Post-Myocardial Infarction Patient Information

Fahrettin Katkat, Çağatay Önal
DAHUDER Medical Journal
Artificial Intelligence in Healthcare and Education
article

Quality and Readability of AI-Generated Post-Myocardial Infarction Patient Information

Fahrettin Katkat, Çağatay Önal
article en

Abstract

Objective: Large language models are increasingly used to obtain health-related information; however, the quality and readability of their responses to practical questions after myocardial infarction (MI) remain uncertain. This study compared ChatGPT and Google Gemini with guideline-based responses for post-MI patient information.Methods: Ten commonly searched questions regarding post-MI recovery and secondary prevention were evaluated. Identical questions were submitted to ChatGPT and Google Gemini, and corresponding reference responses were derived from contemporary European Society of Cardiology guidelines. Two blinded cardiologists independently evaluated response quality using the CLEAR Tool. Readability was assessed using the Flesch Reading Ease Score (FRES) and Flesch–Kincaid Grade Level (FKGL).Results: Thirty responses were evaluated. Inter-rater reliability was excellent (ICC=0.916; 95% CI, 0.821–0.960). Total CLEAR scores differed significantly among the three sources (P=0.002). The median total CLEAR score was highest for ChatGPT [24.00 (23.63–24.38)], followed by guideline-based responses [23.25 (23.00–23.88)] and Google Gemini [19.00 (18.00–20.13)]. ChatGPT and guideline-based responses did not differ significantly in total CLEAR score (P=0.164), whereas Google Gemini scored significantly lower than both. Readability also differed significantly among sources (both FRES and FKGL, P<0.001). Google Gemini had the highest FRES [48.8 (41.0–51.6)] and lowest FKGL [11.5 (10.8–12.1)], indicating greater readability.Conclusion: ChatGPT provided post-MI information with overall quality comparable to guideline-based responses, whereas Google Gemini generated more readable but lower-quality responses. These findings suggest a potential trade-off between readability and informational quality and support the use of LLMs as adjuncts rather than substitutes for clinician-led post-MI patient education.

DAHUDER Medical Journal(Advanced Online Publication)
Istanbul University-Cerrahpaşa (TR), Gazi Hastanesi (TR)
Quality Education
Openalex Percentile: Top 15%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.