Enhanced Swin Transformer Incorporating Multi-Head Self-Attention and Feature Refinement for Fetal Trimester Classification
Background Analyzing fetal ultrasound images is important for estimating gestational age and classifying trimesters, which support fetal growth monitoring and clinical decision-making. Ultrasound imaging is widely used due to its safety, affordability, and non-invasive nature. However, manual interpretation is subject to inter-observer variability and requires significant clinical expertise, limiting its consistency, particularly in resource-constrained settings. Methods This study proposes an enhanced Swin Transformer-based framework for automated fetal trimester classification. The architecture employs a Swin Transformer backbone for hierarchical feature extraction, augmented with a multi-head self-attention mechanism and a feature refinement module to improve contextual representation. The training strategy incorporates Sharpness-Aware Minimization (SAM), MixUp and CutMix data augmentation, a hybrid loss function combining label smoothing, cross-entropy, and focal loss, and test-time augmentation to improve robustness. Results The proposed framework achieves classification accuracies of 98.26% (HC dataset), 97.16% (FL dataset), and 93.53% (HC18 dataset). These represent numerical improvements of 2.0–2.8% over the baseline Swin Transformer. However, pairwise statistical comparisons using McNemar’s test did not reach significance ( p>0.05 for all comparisons). A post-hoc power analysis indicates that the available test set sizes ( n=204−287 ) are sufficient to detect only large performance differences ( ≥ 5–6%), whereas the observed improvements are smaller. The model demonstrates stable performance across 5-fold cross-validation (HC: 97.76% ± 0.98; FL: 97.01% ± 0.62; HC18: 92.71% ± 2.08). Conclusion The proposed approach shows numerically improved performance on internal datasets for fetal trimester classification, but these improvements did not reach statistical significance given the current sample sizes. The framework demonstrates technical feasibility and stable cross-validation performance. Validation on larger, multi-centre datasets ( n≥1,800 ) is required to determine whether the observed improvements generalize and achieve statistical significance. This study highlights the potential of attention-enhanced transformer models while emphasizing the need for adequately powered validation studies.
Authors
- Rashmi Siddalingappa (ORCID: https://orcid.org/0000-0001-9786-8436)
- Shivanand S. Gornale (ORCID: https://orcid.org/0000-0001-5373-4049)
- Prakash S. Hiremath (ORCID: https://orcid.org/0000-0001-7640-6937)
- Priyanka Kamat (ORCID: https://orcid.org/0000-0002-1234-1043)
- Khang Wen Goh
Institutions
- York St John University (GB)
- INTI International University (MY)
- KLE Technological University (IN)
- Rani Channamma University, Belagavi (IN)
Publication Details
- Journal
- F1000Research
- Published
- 2026-10-05
- DOI
- https://doi.org/10.12688/f1000research.186643.1
- Primary Topic
- Medical Image Segmentation Techniques
- Type
- article
- Field-Weighted Citation Impact
- 0.00