Dual-stream audio–visual Transformer fusion for harmful video content detection
The rapid growth of short-form video platforms has increased young audiences’ exposure to harmful content, posing challenges for automated moderation. Existing approaches rely solely on visual information, or incorporate audio only via hand-crafted features or ASR transcripts, overlooking the temporal richness of raw audio waveforms. We propose a dual-stream architecture pairing a PE Core Vision Transformer (ViT) for visual encoding with Wav2Vec2 for raw-waveform acoustic encoding, a combination not previously investigated for harmful video detection on child-directed short-form platforms. The embeddings are fused via a dedicated Transformer module with a learned classification token [CLS]. The architecture is evaluated on three public benchmarks spanning complementary domains: TikHarm (TikTok, 4 classes), MMOB (cartoons, binary), and SAFEPLAY (gameplay, binary and 3-class). Our system achieves 91.05% macro F1 on TikHarm, surpassing the published state of the art by 1.60% despite the competing system exploiting OCR and ASR. On MMOB, it reaches 98.95% accuracy and 98.78% macro F1, a 17.46% gain over the CLIP-based baseline. On SAFEPLAY SubTask 2, it achieves the best reported performance, 85.17% macro F1 (3.46% gain). These results show that end-to-end audio–visual fusion with a Transformer module achieves consistent performance across heterogeneous content domains, each fine-tuned independently, without domain-specific architectural adaptation.
Authors
- Wiem Takrouni
- Hamid Tairi (ORCID: https://orcid.org/0000-0002-5445-0037)
- Mohammed El Jattioui (ORCID: https://orcid.org/0009-0001-4563-5806)
Institutions
- Centre National de la Recherche Scientifique (FR)
- Normandie Université (FR)
- Sidi Mohamed Ben Abdellah University (MA)
- Université de Caen Normandie (FR)
Publication Details
- Journal
- Array
- Published
- 2026-09-18
- DOI
- https://doi.org/10.1016/j.array.2026.101254
- Primary Topic
- Video Analysis and Summarization
- Type
- article
- Field-Weighted Citation Impact
- 0.00