Bridging Artificial and Human Intelligence: Comparative Performance of AI Models and Neurologists in Stroke Recognition- A Cross-Sectional Study

<ns5:p>Background Stroke is a major cause of disability and mortality worldwide. Thus, early detection and intervention, along with appropriate triage, are crucial. The development of large language models (LLMs), such as ChatGPT and Gemini, presents new potential for artificial intelligence in healthcare, including clinical decision support. The objective of this study was to evaluate the diagnostic accuracy and quality of triage recommendations from leading LLMs compared to those from board-certified neurologists in patients with suspected acute stroke. Methods This was a cross-sectional study of 200 posts in the Reddit “AskDocs” section related to possible symptoms of stroke. These posts elicited responses from two LLMs, ChatGPT-4 and Gemini, as well as two board-certified neurologists. Two experienced emergency medicine specialists, who were independent of the survey, evaluated responses for three criteria online using 7-point Likert scales: Ease of Understanding, Scientific Adequacy and Overall Satisfaction. The outcome of interest was advising an Emergency Department (ED) visit. Results Neurologists were much more willing to advocate for visiting the ED (58.5%) than ChatGPT (45%) or Gemini (45%) (p &lt; 0.001). Second, AI models often gave “Unable to determine” response (ChatGPT: 11.5%, Gemini: 14.5%), which was not reported by the neurologists (0%). Regarding quality, neurologist responses were rated highest for Ease of Understanding (Median 7) and Overall Satisfaction (Median 7) (p &lt; 0.001 for both comparisons). Notably, the ChatGPT’s adequacy was rated significantly higher (focus group: Median 6 vs. neurologists: Median 5, p &lt; 0.01), although the levels of Gemini’s scientific adequacy did not differ from those of neurologists (Median 5 vs. 5, p = 0.14). Conclusion LLMs showed strong scientific adequacy and good clarity of presentation; however, their conservative triage strategy raises important patient safety concerns in time-sensitive neurological emergencies such as stroke.</ns5:p>

Authors

Institutions

Publication Details

Journal
F1000Research
Published
2026-07-10
DOI
https://doi.org/10.12688/f1000research.181463.1
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Bridging Artificial and Human Intelligence: Comparative Performance of AI Models and Neurologists in Stroke Recognition- A Cross-Sectional Study

Kaleem Basharat, Ahmed Kassem, Yavuz Yiğit, Doaa Sabir et al.
F1000Research
Artificial Intelligence in Healthcare and Education
article

Bridging Artificial and Human Intelligence: Comparative Performance of AI Models and Neurologists in Stroke Recognition- A Cross-Sectional Study

Kaleem Basharat, Ahmed Kassem, Yavuz Yiğit, Doaa Sabir, Kamal Majed
article en

Abstract

<ns5:p>Background Stroke is a major cause of disability and mortality worldwide. Thus, early detection and intervention, along with appropriate triage, are crucial. The development of large language models (LLMs), such as ChatGPT and Gemini, presents new potential for artificial intelligence in healthcare, including clinical decision support. The objective of this study was to evaluate the diagnostic accuracy and quality of triage recommendations from leading LLMs compared to those from board-certified neurologists in patients with suspected acute stroke. Methods This was a cross-sectional study of 200 posts in the Reddit “AskDocs” section related to possible symptoms of stroke. These posts elicited responses from two LLMs, ChatGPT-4 and Gemini, as well as two board-certified neurologists. Two experienced emergency medicine specialists, who were independent of the survey, evaluated responses for three criteria online using 7-point Likert scales: Ease of Understanding, Scientific Adequacy and Overall Satisfaction. The outcome of interest was advising an Emergency Department (ED) visit. Results Neurologists were much more willing to advocate for visiting the ED (58.5%) than ChatGPT (45%) or Gemini (45%) (p < 0.001). Second, AI models often gave “Unable to determine” response (ChatGPT: 11.5%, Gemini: 14.5%), which was not reported by the neurologists (0%). Regarding quality, neurologist responses were rated highest for Ease of Understanding (Median 7) and Overall Satisfaction (Median 7) (p < 0.001 for both comparisons). Notably, the ChatGPT’s adequacy was rated significantly higher (focus group: Median 6 vs. neurologists: Median 5, p < 0.01), although the levels of Gemini’s scientific adequacy did not differ from those of neurologists (Median 5 vs. 5, p = 0.14). Conclusion LLMs showed strong scientific adequacy and good clarity of presentation; however, their conservative triage strategy raises important patient safety concerns in time-sensitive neurological emergencies such as stroke.</ns5:p>

F1000ResearchVol. 15
Hamad Medical Corporation (QA), Qatar University (QA)
Qatar National Library
Good health and well-being
Openalex Percentile: Top 11%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.