A Multidimensional Evaluation of Large Language Model Responses to the 2024 ESC Hypertension Guidelines: A Comparative Study

Large language models (LLMs), including ChatGPT, Gemini, and DeepSeek, are increasingly used in medicine; however, their performance across multiple clinically relevant domains remains incompletely understood. This cross-sectional comparative study evaluated 225 responses generated by three LLMs to 75 clinical questions selected from the 2024 European Society of Cardiology (ESC) Hypertension Guidelines. Responses were independently assessed by two board-certified cardiologists across five predefined domains: accuracy, clinical relevance, completeness, absence of bias and misinformation, and consistency. No statistically significant differences were observed among the three models in accuracy, clinical relevance, completeness, or absence of bias and misinformation (all p > 0.05). A significant difference was identified in response consistency (p = 0.032), with post hoc analysis demonstrating a significant difference between ChatGPT and Gemini. Overall appropriateness scores did not differ significantly among the three LLMs (p = 0.227). These findings suggest that, although overall performance was comparable, response consistency represents an additional dimension that should be considered when evaluating LLMs for guideline-based clinical applications. Future studies incorporating broader clinical scenarios and updated LLM versions are warranted to further define their role in clinical decision support.

Authors

Institutions

Publication Details

Journal
Journal of Cardiovascular Development and Disease
Published
2026-09-24
DOI
https://doi.org/10.3390/jcdd13100481
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

A Multidimensional Evaluation of Large Language Model Responses to the 2024 ESC Hypertension Guidelines: A Comparative Study

Ferit Böyük, Bilal Çuğlan, Kemal Türker Ulutaş, Halime Tanrıverdi et al.
Journal of Cardiovascular Development and Disease
Artificial Intelligence in Healthcare and Education
article

A Multidimensional Evaluation of Large Language Model Responses to the 2024 ESC Hypertension Guidelines: A Comparative Study

Ferit Böyük, Bilal Çuğlan, Kemal Türker Ulutaş, Halime Tanrıverdi, İsmail Polat Canbolat, Harun Akarsu, Emre Özmen, Aysun Karahan Gün
article en

Abstract

Large language models (LLMs), including ChatGPT, Gemini, and DeepSeek, are increasingly used in medicine; however, their performance across multiple clinically relevant domains remains incompletely understood. This cross-sectional comparative study evaluated 225 responses generated by three LLMs to 75 clinical questions selected from the 2024 European Society of Cardiology (ESC) Hypertension Guidelines. Responses were independently assessed by two board-certified cardiologists across five predefined domains: accuracy, clinical relevance, completeness, absence of bias and misinformation, and consistency. No statistically significant differences were observed among the three models in accuracy, clinical relevance, completeness, or absence of bias and misinformation (all p > 0.05). A significant difference was identified in response consistency (p = 0.032), with post hoc analysis demonstrating a significant difference between ChatGPT and Gemini. Overall appropriateness scores did not differ significantly among the three LLMs (p = 0.227). These findings suggest that, although overall performance was comparable, response consistency represents an additional dimension that should be considered when evaluating LLMs for guideline-based clinical applications. Future studies incorporating broader clinical scenarios and updated LLM versions are warranted to further define their role in clinical decision support.

Journal of Cardiovascular Development and DiseaseVol. 13(10)
Sağlık Bilimleri Üniversitesi (TR), İstanbul Kanuni Sultan Süleyman Eğitim ve Araştırma Hastanesi (TR)
Peace, Justice and strong institutions
Openalex Percentile: Top 15%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.