Hierarchical multimodal semantic fusion for few-shot reader action recognition in reading environments
Abstract Understanding reader behavior in real-world reading environments is important for data-driven content services and intelligent publishing analytics. However, most action recognition methods target motion-dominant scenarios and struggle under few-shot conditions to capture the subtle visual cues and hierarchical semantics of reading-related behaviors. To address this problem, we construct the Reader Action Recognition Dataset (RARS), containing 12 action classes and 1,728 video clips, and propose a Hierarchical Semantic Fusion Network (HSFN) for few-shot reader action recognition. HSFN maintains a shared semantic structure spanning the global, patch, and frame levels across hierarchical visual and textual representations, granularity-consistent multimodal fusion, hyperbolic multi-prototype alignment, and dual-path video-to-video and video-to-text matching. Counterfactual supervision and adaptive-margin multi-grained contrastive learning further refine the boundaries between semantically similar actions. Experiments on RARS and the public SSv2-Small benchmark validate the effectiveness of HSFN under 5-way 1-shot and 5-way 5-shot settings. Ablation studies, attention-guided occlusion, and failure-case analysis further support the proposed components and clarify HSFN’s recognition characteristics for semantically similar actions. Overall, although HSFN remains limited when distinctions between semantically similar actions depend on transient temporal cues or small interacting objects, it provides a viable few-shot solution for fine-grained reader action recognition in intelligent reading environments.
Authors
- Xinong En (ORCID: https://orcid.org/0000-0003-2919-7735)
- Yanping Du (ORCID: https://orcid.org/0000-0003-0211-4223)
- Zhaohua Wang (ORCID: https://orcid.org/0000-0002-3246-3494)
- Jiayi Yang
- Zheng Li
- Yirong Luo
- Jiawen Li
- Yuqian Wang
Publication Details
- Journal
- Scientific Reports
- Published
- 2026-09-30
- DOI
- https://doi.org/10.1038/s41598-026-71846-y
- Primary Topic
- Human Pose and Action Recognition
- Type
- article
- Field-Weighted Citation Impact
- 0.00