HiTCA: fusing hierarchical text and contextual audio for accurate VCR

Automated video content rating (VCR) often neglects essential audio information by primarily focusing on text analysis, resulting in compromised assessment comprehensiveness. To address this shortcoming, we introduce HiTCA (hierarchical text and contextual audio), a novel framework that fuses hierarchical text understanding with contextual audio analysis for accurate rating. The proposed framework leverages a text hierarchical attention transformer, specifically enhanced with dynamic chunking and task-adaptive pre-training, for the comprehensive understanding of lengthy subtitles in movie-length videos. To complement the text analysis, the framework incorporates contextual audio analysis, enabling the identification of rating-relevant sound events and the recognition of emotions in Korean speech. This fusion captures non-textual context often missed in subtitles. Evaluated on a test set derived from 404 Korean movies, HiTCA significantly outperforms a text-only baseline, improving weighted F1 Score from 0.60 to 0.79 and quadratic weighted kappa from 0.77 to 0.91. Notably, the integrated contextual audio substantially enhances discrimination between adjacent rating categories (e.g., the 12+ vs. 15+ boundary). These results demonstrate the effectiveness of fusing advanced hierarchical text processing with specialized contextual audio analysis for more accurate and comprehensive VCR.

Authors

Institutions

Publication Details

Journal
EURASIP Journal on Audio Speech and Music Processing
Published
2026-10-01
DOI
https://doi.org/10.1186/s13636-026-00481-2
Primary Topic
Emotion and Mood Recognition
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

HiTCA: fusing hierarchical text and contextual audio for accurate VCR

Junseok Oh, Ji-Hwan Kim, Juhyeong Nam
EURASIP Journal on Audio Speech and Music Processing
Emotion and Mood Recognition
article

HiTCA: fusing hierarchical text and contextual audio for accurate VCR

Junseok Oh, Ji-Hwan Kim, Juhyeong Nam
article en

Abstract

Automated video content rating (VCR) often neglects essential audio information by primarily focusing on text analysis, resulting in compromised assessment comprehensiveness. To address this shortcoming, we introduce HiTCA (hierarchical text and contextual audio), a novel framework that fuses hierarchical text understanding with contextual audio analysis for accurate rating. The proposed framework leverages a text hierarchical attention transformer, specifically enhanced with dynamic chunking and task-adaptive pre-training, for the comprehensive understanding of lengthy subtitles in movie-length videos. To complement the text analysis, the framework incorporates contextual audio analysis, enabling the identification of rating-relevant sound events and the recognition of emotions in Korean speech. This fusion captures non-textual context often missed in subtitles. Evaluated on a test set derived from 404 Korean movies, HiTCA significantly outperforms a text-only baseline, improving weighted F1 Score from 0.60 to 0.79 and quadratic weighted kappa from 0.77 to 0.91. Notably, the integrated contextual audio substantially enhances discrimination between adjacent rating categories (e.g., the 12+ vs. 15+ boundary). These results demonstrate the effectiveness of fusing advanced hierarchical text processing with specialized contextual audio analysis for more accurate and comprehensive VCR.

EURASIP Journal on Audio Speech and Music Processing
Sogang University (KR), LG (United States) (US)
Peace, Justice and strong institutions, Reduced inequalities
Openalex Percentile: Top 8%
Emotion and Mood Recognition
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

HiTCA: fusing hierarchical text and contextual audio for accurate VCR — Junseok Oh, Ji-Hwan Kim, et al. · EURASIP Journal on Audio Speech and Music Processing (2026) | TGRS Research Map | TGRS