Motion‐Aware Cross‐Modal Modulation for Audiovisual Saliency Prediction

ABSTRACT Audiovisual saliency prediction (AVSP) is a task that simulates how humans combine visual features and audio information to predict the salient regions in a scene. The biggest challenge of this task, which has not been fully solved, is how to maximise the correlation between audio information and visual features without using hardware such as microphone arrays. In view of this, this paper proposes a Motion‐Aware AVSP model (MA‐AVSP). This model focuses on functionally decoupling motion information from traditional visual features and uses it as a global and independent cue to make both static visual features and audio information incorporate motion awareness simultaneously. This facilitates a better fusion between the two modalities, resulting in higher cross‐modal correlation. In this way, MA‐AVSP can adjust the weights of visual and audio features in audiovisual saliency and minimise the distance when visual and audio features are embedded. Specifically, this design processes motion features in a motion‐aware module, which acts as an intermediate layer to regulate the correlation between audiovisual features and sound source localisation at a macroscopic level, thereby achieving more accurate sound source localisation. The aim of this study is not merely to enhance the temporal modelling of vision, but rather to use these motion features as a cross‐modal regulatory factor for macroscopic regulation, in order to achieve more accurate saliency prediction. Experimental results on six commonly used public datasets demonstrate that MA‐AVSP achieves consistent and competitive performance compared with selected baselines, validating the effectiveness of the proposed strategy. Beyond benchmark evaluations, MA‐AVSP provides motion‐guided audiovisual perception and spatial attention priors for expert systems operating in dynamic environments.

Authors

Institutions

Publication Details

Journal
Expert Systems
Published
2026-09-15
DOI
https://doi.org/10.1111/exsy.70427
Primary Topic
Multisensory perception and integration
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Motion‐Aware Cross‐Modal Modulation for Audiovisual Saliency Prediction

Liang Wang, Zhao Wang, Lina Zuo, Zihang Wang et al.
Expert Systems
Multisensory perception and integration
article

Motion‐Aware Cross‐Modal Modulation for Audiovisual Saliency Prediction

Liang Wang, Zhao Wang, Lina Zuo, Zihang Wang, Shaokang Zhang
article en

Abstract

ABSTRACT Audiovisual saliency prediction (AVSP) is a task that simulates how humans combine visual features and audio information to predict the salient regions in a scene. The biggest challenge of this task, which has not been fully solved, is how to maximise the correlation between audio information and visual features without using hardware such as microphone arrays. In view of this, this paper proposes a Motion‐Aware AVSP model (MA‐AVSP). This model focuses on functionally decoupling motion information from traditional visual features and uses it as a global and independent cue to make both static visual features and audio information incorporate motion awareness simultaneously. This facilitates a better fusion between the two modalities, resulting in higher cross‐modal correlation. In this way, MA‐AVSP can adjust the weights of visual and audio features in audiovisual saliency and minimise the distance when visual and audio features are embedded. Specifically, this design processes motion features in a motion‐aware module, which acts as an intermediate layer to regulate the correlation between audiovisual features and sound source localisation at a macroscopic level, thereby achieving more accurate sound source localisation. The aim of this study is not merely to enhance the temporal modelling of vision, but rather to use these motion features as a cross‐modal regulatory factor for macroscopic regulation, in order to achieve more accurate saliency prediction. Experimental results on six commonly used public datasets demonstrate that MA‐AVSP achieves consistent and competitive performance compared with selected baselines, validating the effectiveness of the proposed strategy. Beyond benchmark evaluations, MA‐AVSP provides motion‐guided audiovisual perception and spatial attention priors for expert systems operating in dynamic environments.

Expert SystemsVol. 43(10)
TED University (TR), Hebei University (CN)
Openalex Percentile: Top 7%
Multisensory perception and integration
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.