Motion‐Aware Cross‐Modal Modulation for Audiovisual Saliency Prediction
ABSTRACT Audiovisual saliency prediction (AVSP) is a task that simulates how humans combine visual features and audio information to predict the salient regions in a scene. The biggest challenge of this task, which has not been fully solved, is how to maximise the correlation between audio information and visual features without using hardware such as microphone arrays. In view of this, this paper proposes a Motion‐Aware AVSP model (MA‐AVSP). This model focuses on functionally decoupling motion information from traditional visual features and uses it as a global and independent cue to make both static visual features and audio information incorporate motion awareness simultaneously. This facilitates a better fusion between the two modalities, resulting in higher cross‐modal correlation. In this way, MA‐AVSP can adjust the weights of visual and audio features in audiovisual saliency and minimise the distance when visual and audio features are embedded. Specifically, this design processes motion features in a motion‐aware module, which acts as an intermediate layer to regulate the correlation between audiovisual features and sound source localisation at a macroscopic level, thereby achieving more accurate sound source localisation. The aim of this study is not merely to enhance the temporal modelling of vision, but rather to use these motion features as a cross‐modal regulatory factor for macroscopic regulation, in order to achieve more accurate saliency prediction. Experimental results on six commonly used public datasets demonstrate that MA‐AVSP achieves consistent and competitive performance compared with selected baselines, validating the effectiveness of the proposed strategy. Beyond benchmark evaluations, MA‐AVSP provides motion‐guided audiovisual perception and spatial attention priors for expert systems operating in dynamic environments.
Authors
- Liang Wang (ORCID: https://orcid.org/0000-0003-2839-0832)
- Zhao Wang (ORCID: https://orcid.org/0000-0003-0147-6054)
- Lina Zuo
- Zihang Wang (ORCID: https://orcid.org/0009-0007-1273-3112)
- Shaokang Zhang
Institutions
- TED University (TR)
- Hebei University (CN)
Publication Details
- Journal
- Expert Systems
- Published
- 2026-09-15
- DOI
- https://doi.org/10.1111/exsy.70427
- Primary Topic
- Multisensory perception and integration
- Type
- article
- Field-Weighted Citation Impact
- 0.00