Measurement-Oriented Evaluation of Deep Learning Models for Automated Fetal Head Biometry in Prenatal Ultrasound

Background/Objectives: Accurate fetal head biometry, obtained via ultrasound imaging, is fundamental to prenatal assessment, as it enables fetal growth to be monitored and developmental abnormalities to be detected early. However, conventional manual measurements are time-consuming and operator-dependent, as well as being prone to inter-observer variability. This retrospective study proposes investigating whether conventional segmentation metrics accurately reflect downstream fetal biometric measurements. This would be achieved by evaluating three deep learning architectures on expert-selected, standard-plane fetal head ultrasound images. Methods: Three deep learning architectures, namely, U-Net, Residual U-Net and TransUNet were implemented and evaluated under identical conditions using a local clinical ultrasound dataset of 1918 images and masks from 206 pregnancies. Following segmentation, ellipse fitting was applied to both predicted and reference masks to derive head circumference (HC), biparietal diameter (BPD), and occipitofrontal diameter (OFD). Image-specific physical scaling was applied to convert pixel-based measurements into millimeters. The performance of models was assessed using conventional segmentation metrics, including the Dice Similarity Coefficient (DSC), the Intersection over Union (IoU), and the Hausdorff Distance (HD) metrics. This was complemented by a clinically meaningful biometric evaluation based on the mean absolute error (MAE) of the HC, BPD, and OFD measurements. Results: All three architectures achieved high segmentation performance, with DSC values above 0.98 and IoU values above 0.96. Residual U-Net produced the highest DSC and IoU values (0.9819 ± 0.0102 and 0.9647 ± 0.0196, respectively), whereas TransUNet produced the lowest HD value (1.8935 ± 0.8570 mm). For biometric measurements, the Residual U-Net produced the lowest MAE for HC (1.8877 ± 1.7689 mm) and OFD (0.8672 ± 0.8342 mm), while the TransUNet resulted in the lowest BPD MAE (0.7504 ± 0.7045 mm). Pairwise statistical analysis revealed significant differences in DSC between U-Net and Residual U-Net (p = 0.0172), as well as between Residual U-Net and TransUNet (p = 0.0119). However, no significant differences were observed between U-Net and TransUNet (p = 0.4726). For the biometric measurements, significant differences were only observed for the BPD MAE between the U-Net and TransUNet models (p = 0.0385), while the HC and OFD errors were statistically comparable across the model pairs. Conclusions: This study shows that high overlap-based segmentation performance does not necessarily lead to statistically significant or superior biometric measurement accuracy. The results emphasize the importance of using both conventional segmentation metrics and downstream, measurement-oriented evaluations when comparing deep learning models for fetal head biometry. However, as this study involved a retrospective technical evaluation of curated, expert-selected standard-plane images, the findings could not be interpreted as evidence of clinical effectiveness or immediate clinical applicability. Further validation is required using pregnancy-level separation, multiple observers, independent clinical measurements, pathological and technically challenging examinations, and multicenter external datasets.

Authors

Institutions

Publication Details

Journal
Journal of Clinical Medicine
Published
2026-09-14
DOI
https://doi.org/10.3390/jcm15187127
Primary Topic
Fetal and Pediatric Neurological Disorders
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Measurement-Oriented Evaluation of Deep Learning Models for Automated Fetal Head Biometry in Prenatal Ultrasound

Behiç Akyüz, Gültekin Adanaş Aydın, Furkan Ertürk Urfalı, Emre Dandıl et al.
Journal of Clinical Medicine
Fetal and Pediatric Neurological Disorders
article

Measurement-Oriented Evaluation of Deep Learning Models for Automated Fetal Head Biometry in Prenatal Ultrasound

Behiç Akyüz, Gültekin Adanaş Aydın, Furkan Ertürk Urfalı, Emre Dandıl, Hilal Gülsüm Turan Özsoy, Ebrar Elest Şen
article en

Abstract

Background/Objectives: Accurate fetal head biometry, obtained via ultrasound imaging, is fundamental to prenatal assessment, as it enables fetal growth to be monitored and developmental abnormalities to be detected early. However, conventional manual measurements are time-consuming and operator-dependent, as well as being prone to inter-observer variability. This retrospective study proposes investigating whether conventional segmentation metrics accurately reflect downstream fetal biometric measurements. This would be achieved by evaluating three deep learning architectures on expert-selected, standard-plane fetal head ultrasound images. Methods: Three deep learning architectures, namely, U-Net, Residual U-Net and TransUNet were implemented and evaluated under identical conditions using a local clinical ultrasound dataset of 1918 images and masks from 206 pregnancies. Following segmentation, ellipse fitting was applied to both predicted and reference masks to derive head circumference (HC), biparietal diameter (BPD), and occipitofrontal diameter (OFD). Image-specific physical scaling was applied to convert pixel-based measurements into millimeters. The performance of models was assessed using conventional segmentation metrics, including the Dice Similarity Coefficient (DSC), the Intersection over Union (IoU), and the Hausdorff Distance (HD) metrics. This was complemented by a clinically meaningful biometric evaluation based on the mean absolute error (MAE) of the HC, BPD, and OFD measurements. Results: All three architectures achieved high segmentation performance, with DSC values above 0.98 and IoU values above 0.96. Residual U-Net produced the highest DSC and IoU values (0.9819 ± 0.0102 and 0.9647 ± 0.0196, respectively), whereas TransUNet produced the lowest HD value (1.8935 ± 0.8570 mm). For biometric measurements, the Residual U-Net produced the lowest MAE for HC (1.8877 ± 1.7689 mm) and OFD (0.8672 ± 0.8342 mm), while the TransUNet resulted in the lowest BPD MAE (0.7504 ± 0.7045 mm). Pairwise statistical analysis revealed significant differences in DSC between U-Net and Residual U-Net (p = 0.0172), as well as between Residual U-Net and TransUNet (p = 0.0119). However, no significant differences were observed between U-Net and TransUNet (p = 0.4726). For the biometric measurements, significant differences were only observed for the BPD MAE between the U-Net and TransUNet models (p = 0.0385), while the HC and OFD errors were statistically comparable across the model pairs. Conclusions: This study shows that high overlap-based segmentation performance does not necessarily lead to statistically significant or superior biometric measurement accuracy. The results emphasize the importance of using both conventional segmentation metrics and downstream, measurement-oriented evaluations when comparing deep learning models for fetal head biometry. However, as this study involved a retrospective technical evaluation of curated, expert-selected standard-plane images, the findings could not be interpreted as evidence of clinical effectiveness or immediate clinical applicability. Further validation is required using pregnancy-level separation, multiple observers, independent clinical measurements, pathological and technically challenging examinations, and multicenter external datasets.

Journal of Clinical MedicineVol. 15(18)
Bilecik Şeyh Edebali Üniversitesi (TR), Bursa Technical University (TR), Bursa Yuksek Ihtisas Egitim Ve Arastirma Hastanesi (TR)
Openalex Percentile: Top 7%
Fetal and Pediatric Neurological Disorders
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.