Depth Adaption SegNet for RGB-T Segmentation

RGBT segmentation is a challenging task in the area of computer vision. Current advanced networks for RGBT segmentation focus on extracting deeper discriminative features from RGB and thermal images to provide richer semantic information for the fusion features to the decoder. However, excessively mining deeper semantic features only makes the model redundant. Simultaneously, lacking shallow spatial features leads to difficulties in guaranteeing accurate localization of targets. We believe that the features provided by images can be categorized into three types: edge, patch, and semantics. Only by synchronously taking into account the extraction of all three types of features can models achieve accurate classification on the basis of precise localization. Therefore, we propose Depth Adaption SegNet for RGB-T Segmentation (DASNet). According to the characteristics of the three types of features, we extract semantics, patch, and edge features from the deep, middle, and shallow stages respectively. We specifically design the cross-attention semantics module, patch activation module, and edge enhancement module to perform feature extraction. In addition, in order to efficiently fuse features from different categories, we design a deep-emphasis fusion module to fuse the output features of the modules. Compared to advanced methods, qualitative and quantitative experiments show that DASNet exhibits state-of-the-art performance on the CNN-based RGBT segmentation task.

Authors

Institutions

Publication Details

Journal
Applied Sciences
Published
2026-09-30
DOI
https://doi.org/10.3390/app16199702
Primary Topic
Advanced Neural Network Applications
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Depth Adaption SegNet for RGB-T Segmentation

Shaochuan Zhao, Yong Zhou, Hao Zhu, Bing Liu et al.
Applied Sciences
Advanced Neural Network Applications
article

Depth Adaption SegNet for RGB-T Segmentation

Shaochuan Zhao, Yong Zhou, Hao Zhu, Bing Liu, Chi Zhang
article en

Abstract

RGBT segmentation is a challenging task in the area of computer vision. Current advanced networks for RGBT segmentation focus on extracting deeper discriminative features from RGB and thermal images to provide richer semantic information for the fusion features to the decoder. However, excessively mining deeper semantic features only makes the model redundant. Simultaneously, lacking shallow spatial features leads to difficulties in guaranteeing accurate localization of targets. We believe that the features provided by images can be categorized into three types: edge, patch, and semantics. Only by synchronously taking into account the extraction of all three types of features can models achieve accurate classification on the basis of precise localization. Therefore, we propose Depth Adaption SegNet for RGB-T Segmentation (DASNet). According to the characteristics of the three types of features, we extract semantics, patch, and edge features from the deep, middle, and shallow stages respectively. We specifically design the cross-attention semantics module, patch activation module, and edge enhancement module to perform feature extraction. In addition, in order to efficiently fuse features from different categories, we design a deep-emphasis fusion module to fuse the output features of the modules. Compared to advanced methods, qualitative and quantitative experiments show that DASNet exhibits state-of-the-art performance on the CNN-based RGBT segmentation task.

Applied SciencesVol. 16(19)
China University of Mining and Technology (CN)
Reduced inequalities
Openalex Percentile: Top 14%
Advanced Neural Network Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.