HiTCA: fusing hierarchical text and contextual audio for accurate VCR
Automated video content rating (VCR) often neglects essential audio information by primarily focusing on text analysis, resulting in compromised assessment comprehensiveness. To address this shortcoming, we introduce HiTCA (hierarchical text and contextual audio), a novel framework that fuses hierarchical text understanding with contextual audio analysis for accurate rating. The proposed framework leverages a text hierarchical attention transformer, specifically enhanced with dynamic chunking and task-adaptive pre-training, for the comprehensive understanding of lengthy subtitles in movie-length videos. To complement the text analysis, the framework incorporates contextual audio analysis, enabling the identification of rating-relevant sound events and the recognition of emotions in Korean speech. This fusion captures non-textual context often missed in subtitles. Evaluated on a test set derived from 404 Korean movies, HiTCA significantly outperforms a text-only baseline, improving weighted F1 Score from 0.60 to 0.79 and quadratic weighted kappa from 0.77 to 0.91. Notably, the integrated contextual audio substantially enhances discrimination between adjacent rating categories (e.g., the 12+ vs. 15+ boundary). These results demonstrate the effectiveness of fusing advanced hierarchical text processing with specialized contextual audio analysis for more accurate and comprehensive VCR.
Authors
- Junseok Oh
- Ji-Hwan Kim
- Juhyeong Nam
Institutions
- Sogang University (KR)
- LG (United States) (US)
Publication Details
- Journal
- EURASIP Journal on Audio Speech and Music Processing
- Published
- 2026-10-01
- DOI
- https://doi.org/10.1186/s13636-026-00481-2
- Primary Topic
- Emotion and Mood Recognition
- Type
- article
- Field-Weighted Citation Impact
- 0.00