Severity Stratification Changes What Hallucination Rates Mean: An Item-Matched Audit of Five Large Language Models in Orthodontic Decision Support
Background/Objectives: Large language models (LLMs) are consulted for clinical decision support, and their reliability is judged by hallucination prevalence—a measure treating every unsupported element as equivalent, although a fabricated citation and a fabricated protocol differ in what a clinician acting on them would do. We tested whether prevalence tracks clinical risk. Methods: Five LLMs answered a 100-item orthodontic benchmark validated by three external orthodontists (content validity index 0.923). The reference standard was fixed by construction for the 45 items naming a non-existent entity, so any substantive elaboration is unsupported by design; citations were adjudicated against PubMed and CrossRef. Responses were coded with a seven-category taxonomy and assigned to clinical, operational, or epistemic severity tiers in a post hoc exploratory stratification, pre-specified rather than prospectively registered. Because all models answered the same items, comparisons used Cochran’s Q with pairwise McNemar tests and generalised estimating equations clustered on item. Results: Of 500 responses, 449 (89.8%) contained a hallucination but only 65 (13.0%; 95% CI 10.3–16.2) were clinically consequential: prevalence was roughly sevenfold greater than the rate of clinically consequential output as the authors defined it. Between-model differences were large for undifferentiated prevalence (73–100%; Q = 50.51, p < 0.001) and contracted at the clinical tier (10–16%; Q = 10.00, p = 0.040), where no pairwise contrast survived adjustment. Item-level clustering was far stronger for clinically consequential output than for undifferentiated prevalence (intra-class correlation 0.83 versus 0.07). Conclusions: Prevalence and clinically consequential error are not interchangeable and rank models differently. The stratification is exploratory and author-defined, and the benchmark stress-tests susceptibility to fabricated premises rather than surveying natural use.
Authors
- Ferdi Allaf (ORCID: https://orcid.org/0009-0005-3767-7415)
- Mustafa Özcan (ORCID: https://orcid.org/0009-0008-8331-9493)
Publication Details
- Journal
- Diagnostics
- Published
- 2026-09-30
- DOI
- https://doi.org/10.3390/diagnostics16193180
- Primary Topic
- Artificial Intelligence in Healthcare and Education
- Type
- article
- Field-Weighted Citation Impact
- 0.00