Development and internal validation of a deep learning system for criterion-specific assessment of preclinical dental soap carvings: a multi-rater study

Preclinical dental morphology education, assessment of soap-carving exercises may be influenced by inter-rater variability and the workload associated with manual evaluation. This study developed and internally evaluated an artificial intelligence (AI)–based deep learning system for criterion-specific assessment of soap carvings and examined its agreement with an averaged multi-rater reference standard. A total of 280 soap-carved maxillary left permanent central incisors (Fédération Dentaire Internationale [FDI] tooth 21) were photographed in five standardized views (1,400 images). Three educators independently scored each specimen on ten morphological criteria using a 0–10 scale. Inter-rater reliability was assessed using intraclass correlation coefficients (ICCs), and mean ratings formed criterion-specific averaged multi-rater reference scores. Specimens were partitioned into a development set ( n = 224) and an internal hold-out test set ( n = 56). A ResNet50-based multi-output regression model was trained using five-fold cross-validation across three random seeds and two-stage transfer learning. Performance was evaluated using mean absolute error (MAE), root mean squared error (RMSE), coefficient of determination (R²), Spearman’s rank correlation, ICC, Bland–Altman analysis, and bootstrap 95% confidence intervals (CIs). Gradient-weighted Class Activation Mapping (Grad-CAM) was qualitatively examined in eight hold-out specimens. Average-measure ICCs among educators ranged from 0.849 to 0.994 across the ten criteria. On the hold-out test set, the ensemble achieved a macro MAE of 1.204 (95% CI, 1.059–1.374), macro RMSE of 1.584 (95% CI, 1.371–1.799), and macro R² of 0.223 (95% CI, 0.023–0.348). Criterion-specific Spearman correlations ranged from 0.297 to 0.560, while absolute-agreement ICCs between AI predictions and averaged multi-rater reference scores ranged from 0.309 to 0.429. Criterion 10 had the lowest MAE (0.709), whereas Criterion 9 had the highest (2.029). Bland–Altman mean biases were small, with variable limits of agreement. Grad-CAM showed heterogeneous activation patterns. The findings demonstrate the feasibility of a ResNet50-based approach for criterion-specific assessment of soap carvings, while highlighting substantial variability in criterion-level performance. Absolute agreement with the multi-rater reference standard was heterogeneous and does not support interchangeability with educators assessment or autonomous summative grading. The approach may support human-supervised formative assessment, though prospective external validation is required before implementation.

Authors

Institutions

Publication Details

Journal
BMC Medical Education
Published
2026-10-07
DOI
https://doi.org/10.1186/s12909-026-10554-7
Primary Topic
Dental Radiography and Imaging
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Development and internal validation of a deep learning system for criterion-specific assessment of preclinical dental soap carvings: a multi-rater study

Aysun Sezer, Göknur ÖZTÜRK, Özlem Kara, Mehmet Berk Birdal et al.
BMC Medical Education
Dental Radiography and Imaging
article

Development and internal validation of a deep learning system for criterion-specific assessment of preclinical dental soap carvings: a multi-rater study

Aysun Sezer, Göknur ÖZTÜRK, Özlem Kara, Mehmet Berk Birdal, Hasan Efe Purmut
article en

Abstract

Preclinical dental morphology education, assessment of soap-carving exercises may be influenced by inter-rater variability and the workload associated with manual evaluation. This study developed and internally evaluated an artificial intelligence (AI)–based deep learning system for criterion-specific assessment of soap carvings and examined its agreement with an averaged multi-rater reference standard. A total of 280 soap-carved maxillary left permanent central incisors (Fédération Dentaire Internationale [FDI] tooth 21) were photographed in five standardized views (1,400 images). Three educators independently scored each specimen on ten morphological criteria using a 0–10 scale. Inter-rater reliability was assessed using intraclass correlation coefficients (ICCs), and mean ratings formed criterion-specific averaged multi-rater reference scores. Specimens were partitioned into a development set ( n = 224) and an internal hold-out test set ( n = 56). A ResNet50-based multi-output regression model was trained using five-fold cross-validation across three random seeds and two-stage transfer learning. Performance was evaluated using mean absolute error (MAE), root mean squared error (RMSE), coefficient of determination (R²), Spearman’s rank correlation, ICC, Bland–Altman analysis, and bootstrap 95% confidence intervals (CIs). Gradient-weighted Class Activation Mapping (Grad-CAM) was qualitatively examined in eight hold-out specimens. Average-measure ICCs among educators ranged from 0.849 to 0.994 across the ten criteria. On the hold-out test set, the ensemble achieved a macro MAE of 1.204 (95% CI, 1.059–1.374), macro RMSE of 1.584 (95% CI, 1.371–1.799), and macro R² of 0.223 (95% CI, 0.023–0.348). Criterion-specific Spearman correlations ranged from 0.297 to 0.560, while absolute-agreement ICCs between AI predictions and averaged multi-rater reference scores ranged from 0.309 to 0.429. Criterion 10 had the lowest MAE (0.709), whereas Criterion 9 had the highest (2.029). Bland–Altman mean biases were small, with variable limits of agreement. Grad-CAM showed heterogeneous activation patterns. The findings demonstrate the feasibility of a ResNet50-based approach for criterion-specific assessment of soap carvings, while highlighting substantial variability in criterion-level performance. Absolute agreement with the multi-rater reference standard was heterogeneous and does not support interchangeability with educators assessment or autonomous summative grading. The approach may support human-supervised formative assessment, though prospective external validation is required before implementation.

BMC Medical Education
Biruni University (TR)
Openalex Percentile: Top 9%
Dental Radiography and Imaging
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.