Deep Learning for Schatzker Classification on Anteroposterior Radiographs: A Controlled Benchmark and a Transferable Control Protocol
Background/Objectives: Schatzker type is assigned early, usually from the anteroposterior (AP) radiograph. A single benchmark accuracy cannot say whether a model read the fracture, the anatomy around it, the annotation, or how the archive was assembled. We ran four inexpensive controls to separate those contributions. Methods: We benchmarked a ResNet-50 on PlaTiF, a 2026 public release built for artificial-intelligence research that pairs 421 AP knee radiographs from 186 patients with expert Schatzker labels and per-image tibial segmentations. Evaluation used stratified group five-fold cross-validation grouped by patient, five seeds and balanced accuracy. Inputs were cropped to the expert tibial segmentation shipped with the dataset, an oracle localisation unavailable at deployment. Four controls ran on identical folds: a regression given no pixel content; ablation of the tibial pixels with its complement; a regression on the expert mask alone; and an augmentation audit for label-erasing invariances. Results: Among the 128 fracture patients the network reached 0.345 ± 0.030 six-class balanced accuracy, +0.168 over a non-anatomical baseline fitted on the same folds and the same labels (95% CI +0.106 to +0.230, p = 0.002). Recall was graded: 0.72 for Schatzker VI, 0.11 for V and 0.04 for IV, the last two below chance (0.167). Erasing the tibial pixels left 0.257 ± 0.025, read on its own as the target bone being unused; its complement, the tibia with everything else removed, reached 0.367 ± 0.012, and the whole radiograph, which carries both, only 0.297 ± 0.032 (+0.071 for the tibia alone, 95% CI +0.031 to +0.111, p = 0.008). A regression on the expert mask alone reached 0.213 ± 0.034 and was not distinguishable from the erased model. On fracture versus no classifiable fracture the network reached 0.833 ± 0.028 against 0.814 ± 0.016 for a model given no pixels (p = 0.264), and a coronal computed tomography section accompanied 126 of 128 fracture patients but 24 of 58 others (p = 2.9 × 10−19). Conclusions: Each headline number admitted an explanation other than the fracture in the target bone. An ablation reported without its complement misstated where the signal lay, and an augmentation audit overturned our own explanation for the failure of type IV. Controls of this kind cost minutes, and this study illustrates why they can be informative when a benchmark is built on a retrospective clinical archive.
Authors
- Sohyun Ahn (ORCID: https://orcid.org/0000-0002-0116-3325)
- Sang Hyun Na
Institutions
- Ewha Womans University (KR)
- Ewha Womans University Medical Center (KR)
- Rice University (US)
Publication Details
- Journal
- Journal of Clinical Medicine
- Published
- 2026-09-11
- DOI
- https://doi.org/10.3390/jcm15187075
- Primary Topic
- Total Knee Arthroplasty Outcomes
- Type
- article
- Field-Weighted Citation Impact
- 0.00