Development and internal validation of a deep learning system for criterion-specific assessment of preclinical dental soap carvings: a multi-rater study
Preclinical dental morphology education, assessment of soap-carving exercises may be influenced by inter-rater variability and the workload associated with manual evaluation. This study developed and internally evaluated an artificial intelligence (AI)–based deep learning system for criterion-specific assessment of soap carvings and examined its agreement with an averaged multi-rater reference standard. A total of 280 soap-carved maxillary left permanent central incisors (Fédération Dentaire Internationale [FDI] tooth 21) were photographed in five standardized views (1,400 images). Three educators independently scored each specimen on ten morphological criteria using a 0–10 scale. Inter-rater reliability was assessed using intraclass correlation coefficients (ICCs), and mean ratings formed criterion-specific averaged multi-rater reference scores. Specimens were partitioned into a development set ( n = 224) and an internal hold-out test set ( n = 56). A ResNet50-based multi-output regression model was trained using five-fold cross-validation across three random seeds and two-stage transfer learning. Performance was evaluated using mean absolute error (MAE), root mean squared error (RMSE), coefficient of determination (R²), Spearman’s rank correlation, ICC, Bland–Altman analysis, and bootstrap 95% confidence intervals (CIs). Gradient-weighted Class Activation Mapping (Grad-CAM) was qualitatively examined in eight hold-out specimens. Average-measure ICCs among educators ranged from 0.849 to 0.994 across the ten criteria. On the hold-out test set, the ensemble achieved a macro MAE of 1.204 (95% CI, 1.059–1.374), macro RMSE of 1.584 (95% CI, 1.371–1.799), and macro R² of 0.223 (95% CI, 0.023–0.348). Criterion-specific Spearman correlations ranged from 0.297 to 0.560, while absolute-agreement ICCs between AI predictions and averaged multi-rater reference scores ranged from 0.309 to 0.429. Criterion 10 had the lowest MAE (0.709), whereas Criterion 9 had the highest (2.029). Bland–Altman mean biases were small, with variable limits of agreement. Grad-CAM showed heterogeneous activation patterns. The findings demonstrate the feasibility of a ResNet50-based approach for criterion-specific assessment of soap carvings, while highlighting substantial variability in criterion-level performance. Absolute agreement with the multi-rater reference standard was heterogeneous and does not support interchangeability with educators assessment or autonomous summative grading. The approach may support human-supervised formative assessment, though prospective external validation is required before implementation.
Authors
- Aysun Sezer (ORCID: https://orcid.org/0000-0003-1330-6087)
- Göknur ÖZTÜRK (ORCID: https://orcid.org/0000-0001-8937-2211)
- Özlem Kara (ORCID: https://orcid.org/0000-0002-9878-3917)
- Mehmet Berk Birdal
- Hasan Efe Purmut (ORCID: https://orcid.org/0009-0006-8306-0587)
Institutions
- Biruni University (TR)
Publication Details
- Journal
- BMC Medical Education
- Published
- 2026-10-07
- DOI
- https://doi.org/10.1186/s12909-026-10554-7
- Primary Topic
- Dental Radiography and Imaging
- Type
- article
- Field-Weighted Citation Impact
- 0.00