Evaluation of a deep learning-based system for automated assessment of tooth wear severity on intraoral scans
Artificial intelligence (AI) holds great potential in medical diagnosis owing to its remarkable accuracy and efficiency. We previously proposed a deep learning-based tooth wear severity grading system (TWGS) based on intraoral photographs (PT). This study aims to extend its application to intraoral scans (IOS), improve model accuracy through sample expansion and class balancing, and compare the performance of models trained on diverse datasets. Full-arch occlusal/incisal images (268 IOS and 257 PT) were input into the segmentation model (mask region-based convolutional neural network + self-attention + U-Net algorithm) for training. Subsequently, 3195 IOS and 3117 PT individual tooth images output from the segmentation task were incorporated into the classification model (vision transformer + ResNet algorithm), annotated with Tooth Wear Index (TWI), and randomized into training, validation, and test sets in a 4:1:1 ratio. Class imbalance was addressed through targeted data augmentation strategies. The performance of TWGS was evaluated using several metrics: mean average precision (mAP), accuracy, precision, recall, F1-score, and weighted kappa coefficient. Comparative analysis of classification performance was conducted between anterior and posterior tooth segments, as well as maxillary and mandibular arches. Cross-dataset analysis was performed to evaluate the model’s generalization capability. The accuracy and time required for AI analysis was assessed and compared with that of manual diagnosis. For IOS datasets, the segmentation and classification models demonstrated superior performance, achieving an mAP of 0.91 and a mean accuracy of 0.89. IOS demonstrated superior segmentation performance over PT (mAP: 0.91 vs. 0.85), whereas both modalities yielded comparable accuracy in tooth wear classification (IOS: 0.89; PT: 0.88). After class balancing, the model’s ability to recognize TWI-3 significantly improved, with the mean F1-score increased from 0.84 to 0.89 in IOS, and 0.86 to 0.89 in PT. Comparable diagnostic accuracy was observed across tooth positions and arch locations ( P > 0.05 for both comparisons). Cross-dataset validation showed significant declines in segmentation (mAP: 0.78–0.80) and classification (F1-score: 0.61–0.66) accuracies. TWGS achieved superior grading accuracy compared to junior and mid-level clinicians, while dramatically reducing assessment time from 2.70 ± 1.78 s to 0.08 ± 0.01 s ( P < 0.01). TWGS exhibits superior accuracy and efficiency in determining occlusal/incisal wear degree in both IOS and PT datasets. The class balancing technique significantly improves the performance of TWGS in grading severe wear. TWGS represents a promising adjunct for two-dimensional image‑based assessment of tooth wear severity.
Authors
- Xun Sheng (ORCID: https://orcid.org/0000-0002-9003-4022)
- Lingxiao Zhang (ORCID: https://orcid.org/0000-0001-8922-7639)
- Xin-Shu Dong
- Ya-Ning Pang
- Xin-Yu Mao
- Yang Du
- Jian-Guo Tan
- Ming-Yue Liu
Institutions
- Peking University (CN)
- Kunming Medical University (CN)
- Stomatology Hospital (CN)
Publication Details
- Journal
- BMC Oral Health
- Published
- 2026-09-11
- DOI
- https://doi.org/10.1186/s12903-026-09772-8
- Primary Topic
- Dental Erosion and Treatment
- Type
- article
- Field-Weighted Citation Impact
- 0.00