A reproducible statistical evaluation framework for large-sample assessment of AI-generated medical references: cross-platform application of the reference hallucination score

OBJECTIVE: To quantitatively evaluate the bibliographic reliability of AI-generated medical references across multiple chatbot platforms using the Reference Hallucination Score (RHS) and to examine the influence of output format on reference stability. METHODS: In this cross-sectional comparative study, three AI-based chatbots (ChatGPT, Gemini, and Perplexity) were prompted under 30 predefined medical subheadings in two output formats (letter and review article), generating 3,150 references. Reference reliability was assessed using the RHS, which evaluates presence/verifiability, bibliographic accuracy, PMID validity, and topic relevance. Due to non-normal distribution, data were analyzed using Kruskal-Wallis and Mann-Whitney U tests with Bonferroni correction. Effect sizes (η² and r) were calculated to distinguish statistical detectability from practical magnitude. RESULTS: Significant inter-model differences were detected in total RHS scores (H(2) = 24.88, p < 0.001); however, the overall effect size was very small (η² = 0.007). Pairwise comparisons revealed statistically detectable differences between certain models, although effect sizes were consistently small (r = 0.02-0.20). Format-related differences were observed, with longer outputs demonstrating reduced bibliographic stability; however, these differences were numerically modest and partially attenuated after Bonferroni correction. CONCLUSION: AI-based chatbots exhibit measurable bibliographic instability in reference generation, with statistically detectable but practically small differences between models. Beyond cross-platform comparison, this study proposes an instrument-agnostic statistical framework for large-sample evaluation of AI-generated reference reliability. Positioned within applied methodological research rather than AI benchmarking, the framework may serve as a template for future large-scale AI reliability studies. The proposed framework may contribute to methodological standardization in future evaluations of AI-generated scientific outputs.

Authors

Institutions

Publication Details

Journal
BMC Medical Research Methodology
Published
2026-06-19
DOI
https://doi.org/10.1186/s12874-026-02906-0
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

A reproducible statistical evaluation framework for large-sample assessment of AI-generated medical references: cross-platform application of the reference hallucination score

Mustafa Hüseyin Temel, Yakup Erden, Fatih Bagcier, Nurmuhammet Taş
BMC Medical Research Methodology
Artificial Intelligence in Healthcare and Education
article

A reproducible statistical evaluation framework for large-sample assessment of AI-generated medical references: cross-platform application of the reference hallucination score

Mustafa Hüseyin Temel, Yakup Erden, Fatih Bagcier, Nurmuhammet Taş
article en

Abstract

OBJECTIVE: To quantitatively evaluate the bibliographic reliability of AI-generated medical references across multiple chatbot platforms using the Reference Hallucination Score (RHS) and to examine the influence of output format on reference stability. METHODS: In this cross-sectional comparative study, three AI-based chatbots (ChatGPT, Gemini, and Perplexity) were prompted under 30 predefined medical subheadings in two output formats (letter and review article), generating 3,150 references. Reference reliability was assessed using the RHS, which evaluates presence/verifiability, bibliographic accuracy, PMID validity, and topic relevance. Due to non-normal distribution, data were analyzed using Kruskal-Wallis and Mann-Whitney U tests with Bonferroni correction. Effect sizes (η² and r) were calculated to distinguish statistical detectability from practical magnitude. RESULTS: Significant inter-model differences were detected in total RHS scores (H(2) = 24.88, p < 0.001); however, the overall effect size was very small (η² = 0.007). Pairwise comparisons revealed statistically detectable differences between certain models, although effect sizes were consistently small (r = 0.02-0.20). Format-related differences were observed, with longer outputs demonstrating reduced bibliographic stability; however, these differences were numerically modest and partially attenuated after Bonferroni correction. CONCLUSION: AI-based chatbots exhibit measurable bibliographic instability in reference generation, with statistically detectable but practically small differences between models. Beyond cross-platform comparison, this study proposes an instrument-agnostic statistical framework for large-sample evaluation of AI-generated reference reliability. Positioned within applied methodological research rather than AI benchmarking, the framework may serve as a template for future large-scale AI reliability studies. The proposed framework may contribute to methodological standardization in future evaluations of AI-generated scientific outputs.

BMC Medical Research Methodology
University of Health Science (KH), Erzurum Regional Training and Research Hospital (TR), Sağlık Bilimleri Üniversitesi (TR), Universidad CES (CO), Ophthalmology Clinic (DE), Bolu Abant İzzet Baysal University (TR), University of Health Sciences Antigua (AG)
Openalex Percentile: Top 10%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.