Region-scale saliency transition network for RGB-D salient object detection
RGB-D salient object detection aims to accurately localize visually salient regions in complex scenes by exploiting both appearance semantics from RGB images and geometric structures from depth maps, which is important for RGB-D sensing systems and sensor-assisted visual perception. Existing methods usually rely on attention mechanisms and coarse-to-fine multi-scale fusion strategies to enhance cross-modal feature interactions. However, most existing methods estimate saliency over the entire image or feature map, without explicitly considering how saliency changes across different local observation regions. In practice, saliency is scale-dependent rather than a fixed property that remains unchanged across observation scales. Under global observation, multiple regions may simultaneously exhibit strong visual attractiveness, whereas under local observation, the attention center may be redistributed as the contextual range shrinks and eventually concentrate on more discriminative object regions. To address this issue, this paper proposes a region-scale saliency transition network, termed RSSTNet, which reformulates RGB-D salient object detection as a dynamic saliency modeling problem under changing local observation scales. Specifically, RSSTNet first extracts complementary RGB semantic features and depth geometric features, and constructs content-adaptive multi-scale observation regions based on cross-modal local saliency priors. It then models the saliency transition among different observation scales through region-level representations, enabling the network to explicitly capture attention redistribution from global perception to local focus. Finally, local modality reliability estimation is incorporated to generate saliency maps with complete structures and accurate boundaries, especially when depth measurements are affected by sensor noise, missing values, or modality misalignment. Extensive experiments on public RGB-D salient object detection benchmarks demonstrate that the proposed method achieves competitive performance and shows favorable robustness against complex backgrounds, noisy depth maps, modality conflicts, and multiple salient candidates in real-world sensing scenarios.
Authors
- Lingmin Pu
- Lingui Song
Institutions
- Suzhou Chien-Shiung Institute of Technology (CN)
Publication Details
- Journal
- Scientific Reports
- Published
- 2026-09-15
- DOI
- https://doi.org/10.1038/s41598-026-68335-7
- Primary Topic
- Visual Attention and Saliency Detection
- Type
- article
- Field-Weighted Citation Impact
- 0.00