An Explainable Deep Learning Framework for Musculoskeletal Abnormality Detection

Accurate detection of musculoskeletal abnormalities from radiographs remains challenging because abnormal findings can be subtle and radiographic appearance may vary considerably across examinations. This study aimed to develop and systematically evaluate a deep learning framework incorporating Grad-CAM-based visual attribution for automated musculoskeletal abnormality detection using the publicly available Musculoskeletal Radiographs (MURA) dataset. Four transfer learning architectures, namely DenseNet201, EfficientNetV2-S, ConvNeXt Tiny, and Swin Transformer Tiny, were evaluated for musculoskeletal abnormality classification. The framework incorporated Grad-CAM-based attribution and activation-based weak localization to highlight image regions associated with model predictions without requiring bounding-box annotations during training and a sensitivity-oriented decision threshold to prioritize the detection of abnormal examinations. Model performance was assessed using recall, precision, F1-score, area under the receiver operating characteristic curve (AUC), and confusion matrix analysis. A targeted decision-threshold ablation analysis compared the sensitivity-focused threshold of 0.40 with the conventional 0.50 baseline using the same validation-set probabilities. Previously published results were reviewed to provide contextual information rather than establish comparative superiority. EfficientNetV2-S achieved the highest precision (0.925), F1-score (0.892), and AUC (0.956) on the internal validation set among the evaluated architectures. DenseNet201 achieved the highest recall of 0.876. Previously published performance metrics were reviewed to contextualize the results of the present study. However, differences in datasets, anatomical coverage, preprocessing procedures, and evaluation protocols preclude direct conclusions regarding comparative superiority. At the sensitivity-focused threshold of 0.40, EfficientNetV2-S obtained an accuracy of 0.896 and an F1-score of 0.892 on the internal validation set. Lowering the threshold to 0.40 improved sensitivity from 0.754 to 0.861 and decreased false-negative predictions from 1092 to 618 when compared to the traditional threshold of 0.50. Alongside this improvement, false-positive predictions increased from 150 to 307 and specificity decreased from 0.966 to 0.931. Grad-CAM visualizations and activation-derived bounding boxes provided qualitative information about spatial attribution. However, their correspondence with pathological findings was not independently validated using expert annotations or quantitative localization metrics. On the independent held-out test set, EfficientNetV2-S achieved a recall of 0.826, precision of 0.886, F1-score of 0.855, and AUC of 0.927. The independent evaluation demonstrated lower performance than the internal validation results, highlighting the importance of separating model optimization from final performance assessment. The findings demonstrate the feasibility of integrating transfer learning, Grad-CAM-based visual attribution, activation-based weak localization, and sensitivity-oriented decision-making within a unified framework for musculoskeletal abnormality classification. Although the framework demonstrated promising classification performance, the visual attribution results remain exploratory and do not establish pathological localization accuracy or clinical interpretability. Further evaluation using external clinical datasets, expert-annotated pathological regions, and radiologist assessment is required to establish its generalizability, explanation validity, and potential clinical utility.

Authors

Institutions

Publication Details

Journal
Bioengineering
Published
2026-10-09
DOI
https://doi.org/10.3390/bioengineering13101177
Primary Topic
Medical Imaging and Analysis
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

An Explainable Deep Learning Framework for Musculoskeletal Abnormality Detection

Muhammad Awais, Saeed Ur Rehman, Anwar Ali, Muhammad Saad
Bioengineering
Medical Imaging and Analysis
article

An Explainable Deep Learning Framework for Musculoskeletal Abnormality Detection

Muhammad Awais, Saeed Ur Rehman, Anwar Ali, Muhammad Saad
article en

Abstract

Accurate detection of musculoskeletal abnormalities from radiographs remains challenging because abnormal findings can be subtle and radiographic appearance may vary considerably across examinations. This study aimed to develop and systematically evaluate a deep learning framework incorporating Grad-CAM-based visual attribution for automated musculoskeletal abnormality detection using the publicly available Musculoskeletal Radiographs (MURA) dataset. Four transfer learning architectures, namely DenseNet201, EfficientNetV2-S, ConvNeXt Tiny, and Swin Transformer Tiny, were evaluated for musculoskeletal abnormality classification. The framework incorporated Grad-CAM-based attribution and activation-based weak localization to highlight image regions associated with model predictions without requiring bounding-box annotations during training and a sensitivity-oriented decision threshold to prioritize the detection of abnormal examinations. Model performance was assessed using recall, precision, F1-score, area under the receiver operating characteristic curve (AUC), and confusion matrix analysis. A targeted decision-threshold ablation analysis compared the sensitivity-focused threshold of 0.40 with the conventional 0.50 baseline using the same validation-set probabilities. Previously published results were reviewed to provide contextual information rather than establish comparative superiority. EfficientNetV2-S achieved the highest precision (0.925), F1-score (0.892), and AUC (0.956) on the internal validation set among the evaluated architectures. DenseNet201 achieved the highest recall of 0.876. Previously published performance metrics were reviewed to contextualize the results of the present study. However, differences in datasets, anatomical coverage, preprocessing procedures, and evaluation protocols preclude direct conclusions regarding comparative superiority. At the sensitivity-focused threshold of 0.40, EfficientNetV2-S obtained an accuracy of 0.896 and an F1-score of 0.892 on the internal validation set. Lowering the threshold to 0.40 improved sensitivity from 0.754 to 0.861 and decreased false-negative predictions from 1092 to 618 when compared to the traditional threshold of 0.50. Alongside this improvement, false-positive predictions increased from 150 to 307 and specificity decreased from 0.966 to 0.931. Grad-CAM visualizations and activation-derived bounding boxes provided qualitative information about spatial attribution. However, their correspondence with pathological findings was not independently validated using expert annotations or quantitative localization metrics. On the independent held-out test set, EfficientNetV2-S achieved a recall of 0.826, precision of 0.886, F1-score of 0.855, and AUC of 0.927. The independent evaluation demonstrated lower performance than the internal validation results, highlighting the importance of separating model optimization from final performance assessment. The findings demonstrate the feasibility of integrating transfer learning, Grad-CAM-based visual attribution, activation-based weak localization, and sensitivity-oriented decision-making within a unified framework for musculoskeletal abnormality classification. Although the framework demonstrated promising classification performance, the visual attribution results remain exploratory and do not establish pathological localization accuracy or clinical interpretability. Further evaluation using external clinical datasets, expert-annotated pathological regions, and radiologist assessment is required to establish its generalizability, explanation validity, and potential clinical utility.

BioengineeringVol. 13(10)
Qassim University (SA), University of Hull (GB), Swansea University (GB)
Openalex Percentile: Top 24%
Medical Imaging and Analysis
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.