Audio-visual fusion for near-surface wind speed estimation using surveillance cameras

Surveillance camera-based wind speed (WS) estimation provides a promising and cost-effective solution for near-surface wind sensing; however, its performance remains limited in complex urban environments due to the instability and insufficient accuracy of single-modality observations. This study proposes a surveillance audio-visual fusion framework for WS estimation, termed VA-WindNet. VA-WindNet jointly exploits wind-induced visual motion features and acoustic spectral signatures by integrating convolutional neural networks (CNNs) for spatial feature extraction with temporal neural networks for sequential modeling, together with a sequential bidirectional cross-modal interaction mechanism that progressively refines the visual and acoustic temporal representations before feature aggregation. A multi-scene audio-visual wind dataset of approximately 64 hours was constructed, and 48 VA-WindNet variants were developed by combining multiple CNN backbones with different temporal modeling modules for systematic training and evaluation. Results from real-world experiments show that the VA-WindNet model configured with DenseNet169 and a recurrent neural network achieved the best performance with an RMSE of 1.42 m/s. This corresponds to RMSE reductions of approximately 12.9% and 27.2% compared with the best video-only and audio-only configurations, respectively, and approximately 19% compared with two advanced vision-based wind speed estimation models. Under audio- and video-interference conditions, the median RMSEs of VA-WindNet variants were 1.68 m/s and 2.11 m/s, representing reductions of 33.3% and 19.5% relative to the audio-only and video-only baselines, respectively. These results highlight the potential of surveillance camera networks as a low-cost and infrastructure-efficient solution for high-resolution near-surface wind sensing in urban environments, supporting sustainable urban environmental monitoring, fine-scale wind field analysis, and smart-city-oriented observation enhancement as a complement to existing networks.

Authors

Institutions

Publication Details

Journal
GIScience & Remote Sensing
Published
2026-09-15
DOI
https://doi.org/10.1080/15481603.2026.2731700
Primary Topic
Wind and Air Flow Studies
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Audio-visual fusion for near-surface wind speed estimation using surveillance cameras

Xing Wang, Heyueyang Li, Cuiyan Zhang, Maosu Wang
GIScience & Remote Sensing
Wind and Air Flow Studies
article

Audio-visual fusion for near-surface wind speed estimation using surveillance cameras

Xing Wang, Heyueyang Li, Cuiyan Zhang, Maosu Wang
article en

Abstract

Surveillance camera-based wind speed (WS) estimation provides a promising and cost-effective solution for near-surface wind sensing; however, its performance remains limited in complex urban environments due to the instability and insufficient accuracy of single-modality observations. This study proposes a surveillance audio-visual fusion framework for WS estimation, termed VA-WindNet. VA-WindNet jointly exploits wind-induced visual motion features and acoustic spectral signatures by integrating convolutional neural networks (CNNs) for spatial feature extraction with temporal neural networks for sequential modeling, together with a sequential bidirectional cross-modal interaction mechanism that progressively refines the visual and acoustic temporal representations before feature aggregation. A multi-scene audio-visual wind dataset of approximately 64 hours was constructed, and 48 VA-WindNet variants were developed by combining multiple CNN backbones with different temporal modeling modules for systematic training and evaluation. Results from real-world experiments show that the VA-WindNet model configured with DenseNet169 and a recurrent neural network achieved the best performance with an RMSE of 1.42 m/s. This corresponds to RMSE reductions of approximately 12.9% and 27.2% compared with the best video-only and audio-only configurations, respectively, and approximately 19% compared with two advanced vision-based wind speed estimation models. Under audio- and video-interference conditions, the median RMSEs of VA-WindNet variants were 1.68 m/s and 2.11 m/s, representing reductions of 33.3% and 19.5% relative to the audio-only and video-only baselines, respectively. These results highlight the potential of surveillance camera networks as a low-cost and infrastructure-efficient solution for high-resolution near-surface wind sensing in urban environments, supporting sustainable urban environmental monitoring, fine-scale wind field analysis, and smart-city-oriented observation enhancement as a complement to existing networks.

GIScience & Remote SensingVol. 63(1)
Nanjing Institute of Technology (CN)
Sustainable cities and communities
Openalex Percentile: Top 18%
Wind and Air Flow Studies
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.