Region-adaptive multimodal learning with morphology-guided feature modulation for brain tumor segmentation
Segmentation of brain tumors using multimodal MRI is difficult because patients have different appearances and structure of tumors. The existing methods of multimodal feature fusion achieve global feature integration but they do not recognize the unique properties which exist in each tumor subregion. Subregion-specific feature refinement and morphology-based modulation are implemented by the Region-Adaptive Multimodal Learning (RAML) framework with Attention U-Net architecture. Adaptive feature scaling relies on structural descriptors (region size, compactness, and boundary complexity) and a nested anatomical constraint is used to ensure spatial training hierarchy during training. RAML achieves a mean Dice of 0.9007 (WT), 0.8321 (TC), 0.7572 (ET) over three different seeds on the held-out test set, all trained with the same 50 epoch budget, optimizer, split and evaluation code. RAML did not outperform the baselines in terms of the Dice in this evaluation, with the baselines (U-Net, TransUNet, and Swin-UNetR) having a 1.4–5.6% point improvement in the Dice on the three regions, depending on the region and baseline, with Swin-UNetR using the most parameters (4.7–13.2 times more) than RAML with 1.64 M. This is the result directly without the higher, yet unverified numbers in a previous version of this work which were caused by a problem in a data pipeline that has been addressed. No patient overlap was observed when using the official BraTS inter-year patient mapping, so the dataset BraTS2018 could not be used to evaluate cross-dataset generalization in this study. Of all the advantages of RAML, averaged over the three baselines, this is the model size saving, with a 4.7–13.2 fold model size reduction at the cost of 3.7–4.0%-point reduction in accuracy (mean Dice, U-Net, TransUNet, Swin-UNetR). RAML also features the lowest FLOPs per 2D model compared (3.72 GFLOPs against 9.32–12.21 GFLOPs for baselines), and the second lowest peak inference memory, but is not the fastest in terms of wall-clock latency. We present these results, confirming that the accuracy gap, despite being slightly more significant with the 3D baseline, is statistically robust against all 3 2D baselines (Holm-adjusted p < 0.005 throughout), as the base of all the results for the contribution that this work can currently support, in terms of parameter- and FLOP-efficiency, and not state-of-the-art accuracy.
Authors
- Durgesh Nandan (ORCID: https://orcid.org/0000-0002-9762-3559)
- Sushma Parihar
- Naga Maha Lakshmi K.
Institutions
- Symbiosis International University (IN)
- SR University (IN)
Publication Details
- Journal
- Discover Artificial Intelligence
- Published
- 2026-10-05
- DOI
- https://doi.org/10.1007/s44163-026-02348-z
- Primary Topic
- Brain Tumor Detection and Classification
- Type
- article
- Field-Weighted Citation Impact
- 0.00