Staging and Grading Advanced Periodontitis with Large Language Models from Standardized Case Summaries: A Multicenter Concordance and Reproducibility Study

Background/Objectives: To evaluate the concordance of three general-purpose large language models (LLMs), ChatGPT, Claude, and Gemini, with an adjudicated expert consensus classification of Stage, Grade, and Extent in Stage III–IV periodontitis when applying the 2018 American Academy of Periodontology/European Federation of Periodontology (AAP/EFP) rules to clinician-prepared, standardized case summaries. Methods: In this multicenter retrospective concordance study (two centers, 105 patients), two periodontologists independently classified each case from complete charts and radiographs; discordant cases were adjudicated. De-identified summaries containing pre-extracted clinical, radiographic, and modifier variables were presented to each LLM in three independent sessions per case; the modal decision was the primary index result, and single-session performance was also reported. Stage and Grade concordance (Cohen’s/weighted kappa) were co-primary endpoints; Extent and full concordance were secondary. Results: Modal Stage concordance was almost perfect for ChatGPT (κ = 0.98) and Gemini (κ = 0.92) and substantial for Claude (κ = 0.73); single-session Stage κ ranged from 0.73 to 0.98. Grade (weighted κ = 0.97–1.00) and Extent (κ = 0.93–1.00) concordance were almost perfect. Full concordance ranged from 85.7% (Claude) to 99.0% (ChatGPT). Every classifiable modal Stage discordance understated severity: Claude and Gemini assigned Stage II or I to 14 Stage III cases in which attachment loss met the Stage III threshold but radiographic bone loss was below 33%. Conclusions: Given pre-extracted variables and explicit instructions, the LLMs applied the classification rules with substantial to almost perfect concordance; this reflects rule application, not independent diagnosis. Performance was model-specific, Gemini showed greater session-to-session variability, and two models systematically understaged borderline Stage III cases. The findings support only clinician-supervised, educational use with local validation.

Authors

Institutions

Publication Details

Journal
Journal of Clinical Medicine
Published
2026-09-30
DOI
https://doi.org/10.3390/jcm15197585
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Staging and Grading Advanced Periodontitis with Large Language Models from Standardized Case Summaries: A Multicenter Concordance and Reproducibility Study

Mehmet Kızıltoprak, Mustafa KARACA, Mustafa Özay Uslu
Journal of Clinical Medicine
Artificial Intelligence in Healthcare and Education
article

Staging and Grading Advanced Periodontitis with Large Language Models from Standardized Case Summaries: A Multicenter Concordance and Reproducibility Study

Mehmet Kızıltoprak, Mustafa KARACA, Mustafa Özay Uslu
article en

Abstract

Background/Objectives: To evaluate the concordance of three general-purpose large language models (LLMs), ChatGPT, Claude, and Gemini, with an adjudicated expert consensus classification of Stage, Grade, and Extent in Stage III–IV periodontitis when applying the 2018 American Academy of Periodontology/European Federation of Periodontology (AAP/EFP) rules to clinician-prepared, standardized case summaries. Methods: In this multicenter retrospective concordance study (two centers, 105 patients), two periodontologists independently classified each case from complete charts and radiographs; discordant cases were adjudicated. De-identified summaries containing pre-extracted clinical, radiographic, and modifier variables were presented to each LLM in three independent sessions per case; the modal decision was the primary index result, and single-session performance was also reported. Stage and Grade concordance (Cohen’s/weighted kappa) were co-primary endpoints; Extent and full concordance were secondary. Results: Modal Stage concordance was almost perfect for ChatGPT (κ = 0.98) and Gemini (κ = 0.92) and substantial for Claude (κ = 0.73); single-session Stage κ ranged from 0.73 to 0.98. Grade (weighted κ = 0.97–1.00) and Extent (κ = 0.93–1.00) concordance were almost perfect. Full concordance ranged from 85.7% (Claude) to 99.0% (ChatGPT). Every classifiable modal Stage discordance understated severity: Claude and Gemini assigned Stage II or I to 14 Stage III cases in which attachment loss met the Stage III threshold but radiographic bone loss was below 33%. Conclusions: Given pre-extracted variables and explicit instructions, the LLMs applied the classification rules with substantial to almost perfect concordance; this reflects rule application, not independent diagnosis. Performance was model-specific, Gemini showed greater session-to-session variability, and two models systematically understaged borderline Stage III cases. The findings support only clinician-supervised, educational use with local validation.

Journal of Clinical MedicineVol. 15(19)
Batman University (TR), Alanya University (TR), Department of Public Health (MM), Burdur Mehmet Akif Ersoy Üniversitesi (TR)
Openalex Percentile: Top 15%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.