Benchmarking Generative AI Chatbots in Preventive Medicine: A Comparative Analysis of Accuracy, Completeness, and Inter-Rater Reliability

Background/Objectives: Artificial intelligence (AI) chatbots are increasingly used for healthcare information retrieval, education, and clinical support; however, concerns remain regarding the accuracy and completeness of AI-generated information in preventive medicine, where evidence-based recommendations, epidemiological interpretation, risk assessment, and public health guidance require reliable responses. Comparative evidence across multiple preventive medicine domains remains limited, particularly using standardized expert-validated reference answers. This study aimed to evaluate and compare the accuracy and completeness of responses generated by five AI chatbots across major preventive medicine domains. Methods: A comparative cross-sectional benchmarking study evaluated ChatGPT, DeepSeek, Perplexity AI, Claude AI, and Gemini Advanced using 15 expert-validated questions covering General Preventive Medicine, Clinical Preventive Medicine, Epidemiology and Biostatistics, Occupational Medicine, and Social and Behavioral Science. Responses were independently evaluated by two investigators using predefined accuracy and completeness scoring frameworks. Between-model comparisons were conducted using mixed-effects models accounting for clustering by question, and inter-rater reliability was assessed using agreement statistics. Results: Claude AI achieved the highest overall accuracy (5.67 ± 0.41) and completeness (2.87 ± 0.35), followed by ChatGPT (accuracy, 5.47 ± 0.66; completeness, 2.67 ± 0.49), whereas Gemini Advanced demonstrated the lowest overall performance. Epidemiology and Biostatistics showed the highest pooled completeness score (2.60; 95% CI: 2.32–2.88). Mixed-effects analysis accounting for clustering by question demonstrated a significant overall difference in accuracy among chatbot models (p < 0.001), while inter-rater agreement was high, supporting the consistency of the scoring procedure. Conclusions: Substantial variability in accuracy and completeness was observed among the evaluated AI chatbots, with Claude AI and ChatGPT demonstrating comparatively stronger performance within this specific question set. Given the limited number of unique questions and the continuously evolving nature of AI models, these findings should be considered exploratory and should not be generalized as definitive evidence of model superiority or clinical utility. Larger studies using more diverse question sets and standardized evaluation frameworks are required to confirm these findings.

Authors

Institutions

Publication Details

Journal
Healthcare
Published
2026-10-06
DOI
https://doi.org/10.3390/healthcare14193330
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Benchmarking Generative AI Chatbots in Preventive Medicine: A Comparative Analysis of Accuracy, Completeness, and Inter-Rater Reliability

Abdullah Abdulmohsen Alsabaani, Nojoud Zoraib Alshahrani, Naif Moshabab Alqahtani, Ali Saad Alshahrani et al.
Healthcare
Artificial Intelligence in Healthcare and Education
article

Benchmarking Generative AI Chatbots in Preventive Medicine: A Comparative Analysis of Accuracy, Completeness, and Inter-Rater Reliability

Abdullah Abdulmohsen Alsabaani, Nojoud Zoraib Alshahrani, Naif Moshabab Alqahtani, Ali Saad Alshahrani, Ali Hassan Almaqsudi, Mazen Abdullah N. Alshahrani
article en

Abstract

Background/Objectives: Artificial intelligence (AI) chatbots are increasingly used for healthcare information retrieval, education, and clinical support; however, concerns remain regarding the accuracy and completeness of AI-generated information in preventive medicine, where evidence-based recommendations, epidemiological interpretation, risk assessment, and public health guidance require reliable responses. Comparative evidence across multiple preventive medicine domains remains limited, particularly using standardized expert-validated reference answers. This study aimed to evaluate and compare the accuracy and completeness of responses generated by five AI chatbots across major preventive medicine domains. Methods: A comparative cross-sectional benchmarking study evaluated ChatGPT, DeepSeek, Perplexity AI, Claude AI, and Gemini Advanced using 15 expert-validated questions covering General Preventive Medicine, Clinical Preventive Medicine, Epidemiology and Biostatistics, Occupational Medicine, and Social and Behavioral Science. Responses were independently evaluated by two investigators using predefined accuracy and completeness scoring frameworks. Between-model comparisons were conducted using mixed-effects models accounting for clustering by question, and inter-rater reliability was assessed using agreement statistics. Results: Claude AI achieved the highest overall accuracy (5.67 ± 0.41) and completeness (2.87 ± 0.35), followed by ChatGPT (accuracy, 5.47 ± 0.66; completeness, 2.67 ± 0.49), whereas Gemini Advanced demonstrated the lowest overall performance. Epidemiology and Biostatistics showed the highest pooled completeness score (2.60; 95% CI: 2.32–2.88). Mixed-effects analysis accounting for clustering by question demonstrated a significant overall difference in accuracy among chatbot models (p < 0.001), while inter-rater agreement was high, supporting the consistency of the scoring procedure. Conclusions: Substantial variability in accuracy and completeness was observed among the evaluated AI chatbots, with Claude AI and ChatGPT demonstrating comparatively stronger performance within this specific question set. Given the limited number of unique questions and the continuously evolving nature of AI models, these findings should be considered exploratory and should not be generalized as definitive evidence of model superiority or clinical utility. Larger studies using more diverse question sets and standardized evaluation frameworks are required to confirm these findings.

HealthcareVol. 14(19)
Saudi Arabian Monetary Authority (SA), Saudi Commission for Health Specialties (SA), Asir Central Hospital (SA), Ministry of Defense, Armed Forces Hospital Southern Region, King Khalid University (SA)
Openalex Percentile: Top 19%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.