Severity Stratification Changes What Hallucination Rates Mean: An Item-Matched Audit of Five Large Language Models in Orthodontic Decision Support

Background/Objectives: Large language models (LLMs) are consulted for clinical decision support, and their reliability is judged by hallucination prevalence—a measure treating every unsupported element as equivalent, although a fabricated citation and a fabricated protocol differ in what a clinician acting on them would do. We tested whether prevalence tracks clinical risk. Methods: Five LLMs answered a 100-item orthodontic benchmark validated by three external orthodontists (content validity index 0.923). The reference standard was fixed by construction for the 45 items naming a non-existent entity, so any substantive elaboration is unsupported by design; citations were adjudicated against PubMed and CrossRef. Responses were coded with a seven-category taxonomy and assigned to clinical, operational, or epistemic severity tiers in a post hoc exploratory stratification, pre-specified rather than prospectively registered. Because all models answered the same items, comparisons used Cochran’s Q with pairwise McNemar tests and generalised estimating equations clustered on item. Results: Of 500 responses, 449 (89.8%) contained a hallucination but only 65 (13.0%; 95% CI 10.3–16.2) were clinically consequential: prevalence was roughly sevenfold greater than the rate of clinically consequential output as the authors defined it. Between-model differences were large for undifferentiated prevalence (73–100%; Q = 50.51, p < 0.001) and contracted at the clinical tier (10–16%; Q = 10.00, p = 0.040), where no pairwise contrast survived adjustment. Item-level clustering was far stronger for clinically consequential output than for undifferentiated prevalence (intra-class correlation 0.83 versus 0.07). Conclusions: Prevalence and clinically consequential error are not interchangeable and rank models differently. The stratification is exploratory and author-defined, and the benchmark stress-tests susceptibility to fabricated premises rather than surveying natural use.

Authors

Publication Details

Journal
Diagnostics
Published
2026-09-30
DOI
https://doi.org/10.3390/diagnostics16193180
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Severity Stratification Changes What Hallucination Rates Mean: An Item-Matched Audit of Five Large Language Models in Orthodontic Decision Support

Ferdi Allaf, Mustafa Özcan
Diagnostics
Artificial Intelligence in Healthcare and Education
article

Severity Stratification Changes What Hallucination Rates Mean: An Item-Matched Audit of Five Large Language Models in Orthodontic Decision Support

Ferdi Allaf, Mustafa Özcan
article en

Abstract

Background/Objectives: Large language models (LLMs) are consulted for clinical decision support, and their reliability is judged by hallucination prevalence—a measure treating every unsupported element as equivalent, although a fabricated citation and a fabricated protocol differ in what a clinician acting on them would do. We tested whether prevalence tracks clinical risk. Methods: Five LLMs answered a 100-item orthodontic benchmark validated by three external orthodontists (content validity index 0.923). The reference standard was fixed by construction for the 45 items naming a non-existent entity, so any substantive elaboration is unsupported by design; citations were adjudicated against PubMed and CrossRef. Responses were coded with a seven-category taxonomy and assigned to clinical, operational, or epistemic severity tiers in a post hoc exploratory stratification, pre-specified rather than prospectively registered. Because all models answered the same items, comparisons used Cochran’s Q with pairwise McNemar tests and generalised estimating equations clustered on item. Results: Of 500 responses, 449 (89.8%) contained a hallucination but only 65 (13.0%; 95% CI 10.3–16.2) were clinically consequential: prevalence was roughly sevenfold greater than the rate of clinically consequential output as the authors defined it. Between-model differences were large for undifferentiated prevalence (73–100%; Q = 50.51, p < 0.001) and contracted at the clinical tier (10–16%; Q = 10.00, p = 0.040), where no pairwise contrast survived adjustment. Item-level clustering was far stronger for clinically consequential output than for undifferentiated prevalence (intra-class correlation 0.83 versus 0.07). Conclusions: Prevalence and clinically consequential error are not interchangeable and rank models differently. The stratification is exploratory and author-defined, and the benchmark stress-tests susceptibility to fabricated premises rather than surveying natural use.

DiagnosticsVol. 16(19)
Peace, Justice and strong institutions
Openalex Percentile: Top 16%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Severity Stratification Changes What Hallucination Rates Mean: An Item-Matched Audit of Five Large Language Models in Orthodontic Decision Support — Ferdi Allaf, Mustafa Özcan · Diagnostics (2026) | TGRS Research Map | TGRS