Region-scale saliency transition network for RGB-D salient object detection

RGB-D salient object detection aims to accurately localize visually salient regions in complex scenes by exploiting both appearance semantics from RGB images and geometric structures from depth maps, which is important for RGB-D sensing systems and sensor-assisted visual perception. Existing methods usually rely on attention mechanisms and coarse-to-fine multi-scale fusion strategies to enhance cross-modal feature interactions. However, most existing methods estimate saliency over the entire image or feature map, without explicitly considering how saliency changes across different local observation regions. In practice, saliency is scale-dependent rather than a fixed property that remains unchanged across observation scales. Under global observation, multiple regions may simultaneously exhibit strong visual attractiveness, whereas under local observation, the attention center may be redistributed as the contextual range shrinks and eventually concentrate on more discriminative object regions. To address this issue, this paper proposes a region-scale saliency transition network, termed RSSTNet, which reformulates RGB-D salient object detection as a dynamic saliency modeling problem under changing local observation scales. Specifically, RSSTNet first extracts complementary RGB semantic features and depth geometric features, and constructs content-adaptive multi-scale observation regions based on cross-modal local saliency priors. It then models the saliency transition among different observation scales through region-level representations, enabling the network to explicitly capture attention redistribution from global perception to local focus. Finally, local modality reliability estimation is incorporated to generate saliency maps with complete structures and accurate boundaries, especially when depth measurements are affected by sensor noise, missing values, or modality misalignment. Extensive experiments on public RGB-D salient object detection benchmarks demonstrate that the proposed method achieves competitive performance and shows favorable robustness against complex backgrounds, noisy depth maps, modality conflicts, and multiple salient candidates in real-world sensing scenarios.

Authors

Institutions

Publication Details

Journal
Scientific Reports
Published
2026-09-15
DOI
https://doi.org/10.1038/s41598-026-68335-7
Primary Topic
Visual Attention and Saliency Detection
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Region-scale saliency transition network for RGB-D salient object detection

Lingmin Pu, Lingui Song
Scientific Reports
Visual Attention and Saliency Detection
article

Region-scale saliency transition network for RGB-D salient object detection

Lingmin Pu, Lingui Song
article en

Abstract

RGB-D salient object detection aims to accurately localize visually salient regions in complex scenes by exploiting both appearance semantics from RGB images and geometric structures from depth maps, which is important for RGB-D sensing systems and sensor-assisted visual perception. Existing methods usually rely on attention mechanisms and coarse-to-fine multi-scale fusion strategies to enhance cross-modal feature interactions. However, most existing methods estimate saliency over the entire image or feature map, without explicitly considering how saliency changes across different local observation regions. In practice, saliency is scale-dependent rather than a fixed property that remains unchanged across observation scales. Under global observation, multiple regions may simultaneously exhibit strong visual attractiveness, whereas under local observation, the attention center may be redistributed as the contextual range shrinks and eventually concentrate on more discriminative object regions. To address this issue, this paper proposes a region-scale saliency transition network, termed RSSTNet, which reformulates RGB-D salient object detection as a dynamic saliency modeling problem under changing local observation scales. Specifically, RSSTNet first extracts complementary RGB semantic features and depth geometric features, and constructs content-adaptive multi-scale observation regions based on cross-modal local saliency priors. It then models the saliency transition among different observation scales through region-level representations, enabling the network to explicitly capture attention redistribution from global perception to local focus. Finally, local modality reliability estimation is incorporated to generate saliency maps with complete structures and accurate boundaries, especially when depth measurements are affected by sensor noise, missing values, or modality misalignment. Extensive experiments on public RGB-D salient object detection benchmarks demonstrate that the proposed method achieves competitive performance and shows favorable robustness against complex backgrounds, noisy depth maps, modality conflicts, and multiple salient candidates in real-world sensing scenarios.

Scientific Reports
Suzhou Chien-Shiung Institute of Technology (CN)
Reduced inequalities
Openalex Percentile: Top 13%
Visual Attention and Saliency Detection
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Region-scale saliency transition network for RGB-D salient object detection — Lingmin Pu, Lingui Song · Scientific Reports (2026) | TGRS Research Map | TGRS