Hierarchical multimodal semantic fusion for few-shot reader action recognition in reading environments

Abstract Understanding reader behavior in real-world reading environments is important for data-driven content services and intelligent publishing analytics. However, most action recognition methods target motion-dominant scenarios and struggle under few-shot conditions to capture the subtle visual cues and hierarchical semantics of reading-related behaviors. To address this problem, we construct the Reader Action Recognition Dataset (RARS), containing 12 action classes and 1,728 video clips, and propose a Hierarchical Semantic Fusion Network (HSFN) for few-shot reader action recognition. HSFN maintains a shared semantic structure spanning the global, patch, and frame levels across hierarchical visual and textual representations, granularity-consistent multimodal fusion, hyperbolic multi-prototype alignment, and dual-path video-to-video and video-to-text matching. Counterfactual supervision and adaptive-margin multi-grained contrastive learning further refine the boundaries between semantically similar actions. Experiments on RARS and the public SSv2-Small benchmark validate the effectiveness of HSFN under 5-way 1-shot and 5-way 5-shot settings. Ablation studies, attention-guided occlusion, and failure-case analysis further support the proposed components and clarify HSFN’s recognition characteristics for semantically similar actions. Overall, although HSFN remains limited when distinctions between semantically similar actions depend on transient temporal cues or small interacting objects, it provides a viable few-shot solution for fine-grained reader action recognition in intelligent reading environments.

Authors

Publication Details

Journal
Scientific Reports
Published
2026-09-30
DOI
https://doi.org/10.1038/s41598-026-71846-y
Primary Topic
Human Pose and Action Recognition
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Hierarchical multimodal semantic fusion for few-shot reader action recognition in reading environments

Xinong En, Yanping Du, Zhaohua Wang, Jiayi Yang et al.
Scientific Reports
Human Pose and Action Recognition
article

Hierarchical multimodal semantic fusion for few-shot reader action recognition in reading environments

Xinong En, Yanping Du, Zhaohua Wang, Jiayi Yang, Zheng Li, Yirong Luo, Jiawen Li, Yuqian Wang
article en

Abstract

Abstract Understanding reader behavior in real-world reading environments is important for data-driven content services and intelligent publishing analytics. However, most action recognition methods target motion-dominant scenarios and struggle under few-shot conditions to capture the subtle visual cues and hierarchical semantics of reading-related behaviors. To address this problem, we construct the Reader Action Recognition Dataset (RARS), containing 12 action classes and 1,728 video clips, and propose a Hierarchical Semantic Fusion Network (HSFN) for few-shot reader action recognition. HSFN maintains a shared semantic structure spanning the global, patch, and frame levels across hierarchical visual and textual representations, granularity-consistent multimodal fusion, hyperbolic multi-prototype alignment, and dual-path video-to-video and video-to-text matching. Counterfactual supervision and adaptive-margin multi-grained contrastive learning further refine the boundaries between semantically similar actions. Experiments on RARS and the public SSv2-Small benchmark validate the effectiveness of HSFN under 5-way 1-shot and 5-way 5-shot settings. Ablation studies, attention-guided occlusion, and failure-case analysis further support the proposed components and clarify HSFN’s recognition characteristics for semantically similar actions. Overall, although HSFN remains limited when distinctions between semantically similar actions depend on transient temporal cues or small interacting objects, it provides a viable few-shot solution for fine-grained reader action recognition in intelligent reading environments.

Scientific Reports
Quality Education
Openalex Percentile: Top 14%
Human Pose and Action Recognition
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.