Benchmarking Generative AI Chatbots in Preventive Medicine: A Comparative Analysis of Accuracy, Completeness, and Inter-Rater Reliability
Background/Objectives: Artificial intelligence (AI) chatbots are increasingly used for healthcare information retrieval, education, and clinical support; however, concerns remain regarding the accuracy and completeness of AI-generated information in preventive medicine, where evidence-based recommendations, epidemiological interpretation, risk assessment, and public health guidance require reliable responses. Comparative evidence across multiple preventive medicine domains remains limited, particularly using standardized expert-validated reference answers. This study aimed to evaluate and compare the accuracy and completeness of responses generated by five AI chatbots across major preventive medicine domains. Methods: A comparative cross-sectional benchmarking study evaluated ChatGPT, DeepSeek, Perplexity AI, Claude AI, and Gemini Advanced using 15 expert-validated questions covering General Preventive Medicine, Clinical Preventive Medicine, Epidemiology and Biostatistics, Occupational Medicine, and Social and Behavioral Science. Responses were independently evaluated by two investigators using predefined accuracy and completeness scoring frameworks. Between-model comparisons were conducted using mixed-effects models accounting for clustering by question, and inter-rater reliability was assessed using agreement statistics. Results: Claude AI achieved the highest overall accuracy (5.67 ± 0.41) and completeness (2.87 ± 0.35), followed by ChatGPT (accuracy, 5.47 ± 0.66; completeness, 2.67 ± 0.49), whereas Gemini Advanced demonstrated the lowest overall performance. Epidemiology and Biostatistics showed the highest pooled completeness score (2.60; 95% CI: 2.32–2.88). Mixed-effects analysis accounting for clustering by question demonstrated a significant overall difference in accuracy among chatbot models (p < 0.001), while inter-rater agreement was high, supporting the consistency of the scoring procedure. Conclusions: Substantial variability in accuracy and completeness was observed among the evaluated AI chatbots, with Claude AI and ChatGPT demonstrating comparatively stronger performance within this specific question set. Given the limited number of unique questions and the continuously evolving nature of AI models, these findings should be considered exploratory and should not be generalized as definitive evidence of model superiority or clinical utility. Larger studies using more diverse question sets and standardized evaluation frameworks are required to confirm these findings.
Authors
- Abdullah Abdulmohsen Alsabaani (ORCID: https://orcid.org/0000-0003-2352-8249)
- Nojoud Zoraib Alshahrani
- Naif Moshabab Alqahtani
- Ali Saad Alshahrani (ORCID: https://orcid.org/0000-0002-6721-7752)
- Ali Hassan Almaqsudi
- Mazen Abdullah N. Alshahrani
Institutions
- Saudi Arabian Monetary Authority (SA)
- Saudi Commission for Health Specialties (SA)
- Asir Central Hospital (SA)
- Ministry of Defense
- Armed Forces Hospital Southern Region
- King Khalid University (SA)
Publication Details
- Journal
- Healthcare
- Published
- 2026-10-06
- DOI
- https://doi.org/10.3390/healthcare14193330
- Primary Topic
- Artificial Intelligence in Healthcare and Education
- Type
- article
- Field-Weighted Citation Impact
- 0.00