Vulnerabilities and Defenses in Audio-Visual Attacks: A Survey From Audio to Multimodal Models

With the rapid development of multimodal large language models (MLLMs) and the cost of large amounts of data and resources, researchers are more inclined to use public datasets and fine-tune open-source MLLMs to achieve excellent performance in audiovisual related tasks. While this trend accelerates progress, it also introduces security risks. Empirical studies demonstrate that the latest MLLMs can be manipulated to generate harmful content through crafted inputs, including adversarial perturbations and malicious queries, bypassing internal safeguards. To gain a deeper comprehension of the inherent security vulnerabilities associated with audio-visual-based multimodal models, a series of surveys investigates various types of attacks, including adversarial and backdoor attacks. While existing surveys on audio-visual attacks provide a comprehensive overview, they are limited to specific types of attacks, which lack a unified review of various types of attacks. To bridge this gap, this paper presents a comprehensive and systematic review of audio-visual attacks, which include adversarial, backdoor, and jailbreak attacks. Furthermore, we also review various types of attacks in the latest audio-visual-based MLLMs, a dimension notably absent in existing surveys. Drawing from extensive literature, this paper delineates both challenges and emergent trends for future research on audio-visual attacks and defenses.

Authors

Institutions

Publication Details

Journal
ACM Computing Surveys
Published
2026-10-05
DOI
https://doi.org/10.1145/3857219
Primary Topic
Adversarial Robustness in Machine Learning
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Vulnerabilities and Defenses in Audio-Visual Attacks: A Survey From Audio to Multimodal Models

Shuai Zhao, Jinming Wen, Yuwen Li, XC Wu et al.
ACM Computing Surveys
Adversarial Robustness in Machine Learning
article

Vulnerabilities and Defenses in Audio-Visual Attacks: A Survey From Audio to Multimodal Models

Shuai Zhao, Jinming Wen, Yuwen Li, XC Wu, Yanhao Jia
article en

Abstract

With the rapid development of multimodal large language models (MLLMs) and the cost of large amounts of data and resources, researchers are more inclined to use public datasets and fine-tune open-source MLLMs to achieve excellent performance in audiovisual related tasks. While this trend accelerates progress, it also introduces security risks. Empirical studies demonstrate that the latest MLLMs can be manipulated to generate harmful content through crafted inputs, including adversarial perturbations and malicious queries, bypassing internal safeguards. To gain a deeper comprehension of the inherent security vulnerabilities associated with audio-visual-based multimodal models, a series of surveys investigates various types of attacks, including adversarial and backdoor attacks. While existing surveys on audio-visual attacks provide a comprehensive overview, they are limited to specific types of attacks, which lack a unified review of various types of attacks. To bridge this gap, this paper presents a comprehensive and systematic review of audio-visual attacks, which include adversarial, backdoor, and jailbreak attacks. Furthermore, we also review various types of attacks in the latest audio-visual-based MLLMs, a dimension notably absent in existing surveys. Drawing from extensive literature, this paper delineates both challenges and emergent trends for future research on audio-visual attacks and defenses.

ACM Computing Surveys
Northeastern University (US), Nanyang Technological University (SG), Shanghai Jiao Tong University (CN), Jilin University (CN)
Openalex Percentile: Top 10%
Adversarial Robustness in Machine Learning
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Vulnerabilities and Defenses in Audio-Visual Attacks: A Survey From Audio to Multimodal Models — Shuai Zhao, Jinming Wen, et al. · ACM Computing Surveys (2026) | TGRS Research Map | TGRS