Cross-Dataset Evaluation of the Lightweight YOLO Family for Breast Ultrasound Lesion Segmentation: Effects of Preprocessing, Hyperparameter Optimization, and Test-Time Augmentation

Breast cancer remains a major global health concern, making accurate lesion assessment essential for effective clinical decision making. Deep learning has shown promising performance in breast ultrasound analysis, yet models evaluated on data from the same source may not generalize reliably to images acquired using different scanners and acquisition settings. This study therefore examines cross-dataset generalization and the factors that can improve it. Five lightweight YOLO instance-segmentation architectures (YOLOv8n, YOLO11n, YOLO11s, YOLO26n, and YOLO26s) were trained using a patient-grouped BUS-BRA protocol and a single training seed, and evaluated internally on held-out data and externally on BUS-UCLM, which served as the single target dataset. Ultrasound-specific preprocessing, test-time augmentation (TTA), and Optuna-selected configurations were assessed across 40 paired internal–external comparisons. Cross-dataset evaluation demonstrated consistent lesion delineation on BUS-UCLM, with matched Dice ranging from 0.861 to 0.875 and matched IoU from 0.764 to 0.786 across the five architectures. Preprocessing improved mask mAP@50–95 across all ten checkpoints on both datasets, while TTA improved detection-adjusted Dice despite having little effect on mAP@50–95. Optuna tuning improved external mAP@50–95 across all five architectures despite limited internal gains. Grad-CAM++ showed predominantly lesion-centered attention across both datasets, while ONNX Runtime deployment achieved 5.65–13.62 FPS on CPU. These findings highlight the importance of external validation, detection-aware evaluation, and efficient deployment for reliable breast ultrasound segmentation.

Authors

Institutions

Publication Details

Journal
NDT
Published
2026-09-16
DOI
https://doi.org/10.3390/ndt4030028
Primary Topic
AI in cancer detection
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Cross-Dataset Evaluation of the Lightweight YOLO Family for Breast Ultrasound Lesion Segmentation: Effects of Preprocessing, Hyperparameter Optimization, and Test-Time Augmentation

Sarnali Basak, Safiul Haque Chowdhury, Pial Ghosh, Rakib Ahammed Diptho
NDT
AI in cancer detection
article

Cross-Dataset Evaluation of the Lightweight YOLO Family for Breast Ultrasound Lesion Segmentation: Effects of Preprocessing, Hyperparameter Optimization, and Test-Time Augmentation

Sarnali Basak, Safiul Haque Chowdhury, Pial Ghosh, Rakib Ahammed Diptho
article en

Abstract

Breast cancer remains a major global health concern, making accurate lesion assessment essential for effective clinical decision making. Deep learning has shown promising performance in breast ultrasound analysis, yet models evaluated on data from the same source may not generalize reliably to images acquired using different scanners and acquisition settings. This study therefore examines cross-dataset generalization and the factors that can improve it. Five lightweight YOLO instance-segmentation architectures (YOLOv8n, YOLO11n, YOLO11s, YOLO26n, and YOLO26s) were trained using a patient-grouped BUS-BRA protocol and a single training seed, and evaluated internally on held-out data and externally on BUS-UCLM, which served as the single target dataset. Ultrasound-specific preprocessing, test-time augmentation (TTA), and Optuna-selected configurations were assessed across 40 paired internal–external comparisons. Cross-dataset evaluation demonstrated consistent lesion delineation on BUS-UCLM, with matched Dice ranging from 0.861 to 0.875 and matched IoU from 0.764 to 0.786 across the five architectures. Preprocessing improved mask mAP@50–95 across all ten checkpoints on both datasets, while TTA improved detection-adjusted Dice despite having little effect on mAP@50–95. Optuna tuning improved external mAP@50–95 across all five architectures despite limited internal gains. Grad-CAM++ showed predominantly lesion-centered attention across both datasets, while ONNX Runtime deployment achieved 5.65–13.62 FPS on CPU. These findings highlight the importance of external validation, detection-aware evaluation, and efficient deployment for reliable breast ultrasound segmentation.

NDTVol. 4(3)
Jahangirnagar University (BD), BRAC University (BD)
Peace, Justice and strong institutions
Openalex Percentile: Top 8%
AI in cancer detection
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.