Dual-Encoder Network Optimization for Binary Semantic Segmentation of Focal Liver Lesions (FLLs)
Automated binary semantic segmentation of focal liver lesions (FLLs) in multi-phasic computed tomography (CT) scans is important for knowledge extraction of hepatic malignancy, staging disease, and planning treatment. However, standard models struggle due to extreme class imbalances, complex boundary transitions, and tissue heterogeneity across phases. We present a systematic ablation study evaluating CNN Only, Transformer Only, and parallel dual-encoder networks that combine a local convolutional neural network (CNN) branch (VGG-19 or ConvNeXt-Tiny) with a global hierarchical vision transformer branch (Swin or CSwin) against standard and state-of-the-art baselines for semantic segmentation in medical images. These backbones were evaluated across feature fusions as: early, intermediate, and late strategies on a 308-patient subset of the MCT-LTDiag dataset, equally partitioned across five FLL subtypes (BCLM, CRLM, HCC, HH, and ICC). The study evaluates binary lesion segmentation (Foreground vs. Background Liver) across 5 distinct clinical subtype cohorts (BCLM, CRLM, HCC, HH, and ICC). Evaluation was conducted via patient-grouped, stratified three-fold cross-validation. Pairwise architectural comparisons were performed using a patient-level paired design, with statistical significance evaluated via bootstrap 95% confidence intervals (10,000 resamples), two-sided paired permutation tests (10,000 sign-flips), and Cohen’s du effect sizes adjusted using the Holm-Bonferroni method. The proposed dual-encoder configuration utilizing ConvNeXt-Tiny + Swin Transformer with intermediate feature fusion achieved an optimal trade-off in segmentation performance, yielding a top Mean Dice score of 0.875 ± 0.076 and a Mean IoU of 0.861 ± 0.082. Compared to the best CNN-only baseline (ConvNeXt-Tiny), the proposed model showed statistically significant improvements across all parameters: +0.032 Mean Dice, −15.00 mm HD95, and +0.137 Normalized Surface Dice (NSD). From a resource-complexity perspective, the optimal ConvNeXt-Tiny + Swin Transformer with Intermediate Fusion model operated at <8.1% of the GFLOPs (33.34 vs. 412.30 GFLOPs) and achieved ~6.7 times faster with inference speedup (24.66 ms vs. 165.00 ms latency) than computationally intensive state-of-the-art 2D nnU-Net baselines.
Authors
- Edwin Sybingco (ORCID: https://orcid.org/0000-0003-1296-3616)
- Justin Ryan L. Tan (ORCID: https://orcid.org/0000-0003-0606-9757)
- Melvin K. Cabatuan (ORCID: https://orcid.org/0000-0002-8864-7233)
- Laurence A. Gan Lim (ORCID: https://orcid.org/0000-0002-5049-7428)
- Jessica S. Velasco (ORCID: https://orcid.org/0000-0002-2164-8522)
- Argel Bandala
- James Manuel Medalla
- Rennan Baldovino
- Stephen Wong
- Cesar Llorente
Institutions
- Chinese General Hospital and Medical Center (PH)
- De La Salle University (PH)
Publication Details
- Journal
- Machine Learning and Knowledge Extraction
- Published
- 2026-10-09
- DOI
- https://doi.org/10.3390/make8100323
- Primary Topic
- Medical Image Segmentation Techniques
- Type
- article
- Field-Weighted Citation Impact
- 0.00