Deep Learning for Schatzker Classification on Anteroposterior Radiographs: A Controlled Benchmark and a Transferable Control Protocol

Background/Objectives: Schatzker type is assigned early, usually from the anteroposterior (AP) radiograph. A single benchmark accuracy cannot say whether a model read the fracture, the anatomy around it, the annotation, or how the archive was assembled. We ran four inexpensive controls to separate those contributions. Methods: We benchmarked a ResNet-50 on PlaTiF, a 2026 public release built for artificial-intelligence research that pairs 421 AP knee radiographs from 186 patients with expert Schatzker labels and per-image tibial segmentations. Evaluation used stratified group five-fold cross-validation grouped by patient, five seeds and balanced accuracy. Inputs were cropped to the expert tibial segmentation shipped with the dataset, an oracle localisation unavailable at deployment. Four controls ran on identical folds: a regression given no pixel content; ablation of the tibial pixels with its complement; a regression on the expert mask alone; and an augmentation audit for label-erasing invariances. Results: Among the 128 fracture patients the network reached 0.345 ± 0.030 six-class balanced accuracy, +0.168 over a non-anatomical baseline fitted on the same folds and the same labels (95% CI +0.106 to +0.230, p = 0.002). Recall was graded: 0.72 for Schatzker VI, 0.11 for V and 0.04 for IV, the last two below chance (0.167). Erasing the tibial pixels left 0.257 ± 0.025, read on its own as the target bone being unused; its complement, the tibia with everything else removed, reached 0.367 ± 0.012, and the whole radiograph, which carries both, only 0.297 ± 0.032 (+0.071 for the tibia alone, 95% CI +0.031 to +0.111, p = 0.008). A regression on the expert mask alone reached 0.213 ± 0.034 and was not distinguishable from the erased model. On fracture versus no classifiable fracture the network reached 0.833 ± 0.028 against 0.814 ± 0.016 for a model given no pixels (p = 0.264), and a coronal computed tomography section accompanied 126 of 128 fracture patients but 24 of 58 others (p = 2.9 × 10−19). Conclusions: Each headline number admitted an explanation other than the fracture in the target bone. An ablation reported without its complement misstated where the signal lay, and an augmentation audit overturned our own explanation for the failure of type IV. Controls of this kind cost minutes, and this study illustrates why they can be informative when a benchmark is built on a retrospective clinical archive.

Authors

Institutions

Publication Details

Journal
Journal of Clinical Medicine
Published
2026-09-11
DOI
https://doi.org/10.3390/jcm15187075
Primary Topic
Total Knee Arthroplasty Outcomes
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Deep Learning for Schatzker Classification on Anteroposterior Radiographs: A Controlled Benchmark and a Transferable Control Protocol

Sohyun Ahn, Sang Hyun Na
Journal of Clinical Medicine
Total Knee Arthroplasty Outcomes
article

Deep Learning for Schatzker Classification on Anteroposterior Radiographs: A Controlled Benchmark and a Transferable Control Protocol

Sohyun Ahn, Sang Hyun Na
article en

Abstract

Background/Objectives: Schatzker type is assigned early, usually from the anteroposterior (AP) radiograph. A single benchmark accuracy cannot say whether a model read the fracture, the anatomy around it, the annotation, or how the archive was assembled. We ran four inexpensive controls to separate those contributions. Methods: We benchmarked a ResNet-50 on PlaTiF, a 2026 public release built for artificial-intelligence research that pairs 421 AP knee radiographs from 186 patients with expert Schatzker labels and per-image tibial segmentations. Evaluation used stratified group five-fold cross-validation grouped by patient, five seeds and balanced accuracy. Inputs were cropped to the expert tibial segmentation shipped with the dataset, an oracle localisation unavailable at deployment. Four controls ran on identical folds: a regression given no pixel content; ablation of the tibial pixels with its complement; a regression on the expert mask alone; and an augmentation audit for label-erasing invariances. Results: Among the 128 fracture patients the network reached 0.345 ± 0.030 six-class balanced accuracy, +0.168 over a non-anatomical baseline fitted on the same folds and the same labels (95% CI +0.106 to +0.230, p = 0.002). Recall was graded: 0.72 for Schatzker VI, 0.11 for V and 0.04 for IV, the last two below chance (0.167). Erasing the tibial pixels left 0.257 ± 0.025, read on its own as the target bone being unused; its complement, the tibia with everything else removed, reached 0.367 ± 0.012, and the whole radiograph, which carries both, only 0.297 ± 0.032 (+0.071 for the tibia alone, 95% CI +0.031 to +0.111, p = 0.008). A regression on the expert mask alone reached 0.213 ± 0.034 and was not distinguishable from the erased model. On fracture versus no classifiable fracture the network reached 0.833 ± 0.028 against 0.814 ± 0.016 for a model given no pixels (p = 0.264), and a coronal computed tomography section accompanied 126 of 128 fracture patients but 24 of 58 others (p = 2.9 × 10−19). Conclusions: Each headline number admitted an explanation other than the fracture in the target bone. An ablation reported without its complement misstated where the signal lay, and an augmentation audit overturned our own explanation for the failure of type IV. Controls of this kind cost minutes, and this study illustrates why they can be informative when a benchmark is built on a retrospective clinical archive.

Journal of Clinical MedicineVol. 15(18)
Ewha Womans University (KR), Ewha Womans University Medical Center (KR), Rice University (US)
Openalex Percentile: Top 8%
Total Knee Arthroplasty Outcomes
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.