Dual-stream audio–visual Transformer fusion for harmful video content detection

The rapid growth of short-form video platforms has increased young audiences’ exposure to harmful content, posing challenges for automated moderation. Existing approaches rely solely on visual information, or incorporate audio only via hand-crafted features or ASR transcripts, overlooking the temporal richness of raw audio waveforms. We propose a dual-stream architecture pairing a PE Core Vision Transformer (ViT) for visual encoding with Wav2Vec2 for raw-waveform acoustic encoding, a combination not previously investigated for harmful video detection on child-directed short-form platforms. The embeddings are fused via a dedicated Transformer module with a learned classification token [CLS]. The architecture is evaluated on three public benchmarks spanning complementary domains: TikHarm (TikTok, 4 classes), MMOB (cartoons, binary), and SAFEPLAY (gameplay, binary and 3-class). Our system achieves 91.05% macro F1 on TikHarm, surpassing the published state of the art by 1.60% despite the competing system exploiting OCR and ASR. On MMOB, it reaches 98.95% accuracy and 98.78% macro F1, a 17.46% gain over the CLIP-based baseline. On SAFEPLAY SubTask 2, it achieves the best reported performance, 85.17% macro F1 (3.46% gain). These results show that end-to-end audio–visual fusion with a Transformer module achieves consistent performance across heterogeneous content domains, each fine-tuned independently, without domain-specific architectural adaptation.

Authors

Institutions

Publication Details

Journal
Array
Published
2026-09-18
DOI
https://doi.org/10.1016/j.array.2026.101254
Primary Topic
Video Analysis and Summarization
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Dual-stream audio–visual Transformer fusion for harmful video content detection

Wiem Takrouni, Hamid Tairi, Mohammed El Jattioui
Array
Video Analysis and Summarization
article

Dual-stream audio–visual Transformer fusion for harmful video content detection

Wiem Takrouni, Hamid Tairi, Mohammed El Jattioui
article en

Abstract

The rapid growth of short-form video platforms has increased young audiences’ exposure to harmful content, posing challenges for automated moderation. Existing approaches rely solely on visual information, or incorporate audio only via hand-crafted features or ASR transcripts, overlooking the temporal richness of raw audio waveforms. We propose a dual-stream architecture pairing a PE Core Vision Transformer (ViT) for visual encoding with Wav2Vec2 for raw-waveform acoustic encoding, a combination not previously investigated for harmful video detection on child-directed short-form platforms. The embeddings are fused via a dedicated Transformer module with a learned classification token [CLS]. The architecture is evaluated on three public benchmarks spanning complementary domains: TikHarm (TikTok, 4 classes), MMOB (cartoons, binary), and SAFEPLAY (gameplay, binary and 3-class). Our system achieves 91.05% macro F1 on TikHarm, surpassing the published state of the art by 1.60% despite the competing system exploiting OCR and ASR. On MMOB, it reaches 98.95% accuracy and 98.78% macro F1, a 17.46% gain over the CLIP-based baseline. On SAFEPLAY SubTask 2, it achieves the best reported performance, 85.17% macro F1 (3.46% gain). These results show that end-to-end audio–visual fusion with a Transformer module achieves consistent performance across heterogeneous content domains, each fine-tuned independently, without domain-specific architectural adaptation.

ArrayVol. 32
Centre National de la Recherche Scientifique (FR), Normandie Université (FR), Sidi Mohamed Ben Abdellah University (MA), Université de Caen Normandie (FR)
Openalex Percentile: Top 13%
Video Analysis and Summarization
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Dual-stream audio–visual Transformer fusion for harmful video content detection — Wiem Takrouni, Hamid Tairi, et al. · Array (2026) | TGRS Research Map | TGRS