Staging and Grading Advanced Periodontitis with Large Language Models from Standardized Case Summaries: A Multicenter Concordance and Reproducibility Study
Background/Objectives: To evaluate the concordance of three general-purpose large language models (LLMs), ChatGPT, Claude, and Gemini, with an adjudicated expert consensus classification of Stage, Grade, and Extent in Stage III–IV periodontitis when applying the 2018 American Academy of Periodontology/European Federation of Periodontology (AAP/EFP) rules to clinician-prepared, standardized case summaries. Methods: In this multicenter retrospective concordance study (two centers, 105 patients), two periodontologists independently classified each case from complete charts and radiographs; discordant cases were adjudicated. De-identified summaries containing pre-extracted clinical, radiographic, and modifier variables were presented to each LLM in three independent sessions per case; the modal decision was the primary index result, and single-session performance was also reported. Stage and Grade concordance (Cohen’s/weighted kappa) were co-primary endpoints; Extent and full concordance were secondary. Results: Modal Stage concordance was almost perfect for ChatGPT (κ = 0.98) and Gemini (κ = 0.92) and substantial for Claude (κ = 0.73); single-session Stage κ ranged from 0.73 to 0.98. Grade (weighted κ = 0.97–1.00) and Extent (κ = 0.93–1.00) concordance were almost perfect. Full concordance ranged from 85.7% (Claude) to 99.0% (ChatGPT). Every classifiable modal Stage discordance understated severity: Claude and Gemini assigned Stage II or I to 14 Stage III cases in which attachment loss met the Stage III threshold but radiographic bone loss was below 33%. Conclusions: Given pre-extracted variables and explicit instructions, the LLMs applied the classification rules with substantial to almost perfect concordance; this reflects rule application, not independent diagnosis. Performance was model-specific, Gemini showed greater session-to-session variability, and two models systematically understaged borderline Stage III cases. The findings support only clinician-supervised, educational use with local validation.
Authors
- Mehmet Kızıltoprak (ORCID: https://orcid.org/0000-0001-5829-4812)
- Mustafa KARACA (ORCID: https://orcid.org/0000-0001-5853-2366)
- Mustafa Özay Uslu (ORCID: https://orcid.org/0000-0002-9707-1379)
Institutions
- Batman University (TR)
- Alanya University (TR)
- Department of Public Health (MM)
- Burdur Mehmet Akif Ersoy Üniversitesi (TR)
Publication Details
- Journal
- Journal of Clinical Medicine
- Published
- 2026-09-30
- DOI
- https://doi.org/10.3390/jcm15197585
- Primary Topic
- Artificial Intelligence in Healthcare and Education
- Type
- article
- Field-Weighted Citation Impact
- 0.00