Application of human action recognition in film and television editing based on multi-scale visual transformer and attention mechanism

Abstract With the rapid growth of the scale of film and television content and the increasing demand for intelligent editing processes, problems such as large time scale spans, strong spatial interference, and weak continuity of action semantics in human actions in film and television works have become increasingly prominent. Therefore, the research aims to construct a human action recognition framework that integrates multi-scale visual Transformer and spatial–temporal attention mechanisms. Multi-scale context encoding is used to describe the long-term and short-term temporal correlation of actions, and multi-scale spatial and temporal attention networks are introduced to explicitly enhance key human body regions and key action fragments. Cross-scale feature fusion and action aggregation are combined to achieve robust discrimination. When the occlusion ratio reached 30%, the model recognition performance remained above 0.86, and the fluctuation amplitude was controlled within 0.02. In the transitional action detection task, the key action hit rate reached the highest 0.96, showing the ability to effectively capture the action switching process. The average detection delay was stable at around 30 ms, which was about 20% lower than that of the comparison method. The proposed multi-scale modeling and attention collaboration strategy can effectively improve the accuracy, stability and robustness of human action recognition in complex film and television scenes, providing a feasible technical path for intelligent film and television editing and content understanding.

Authors

Institutions

Publication Details

Journal
Discover Artificial Intelligence
Published
2026-09-05
DOI
https://doi.org/10.1007/s44163-026-01904-x
Primary Topic
Human Pose and Action Recognition
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Application of human action recognition in film and television editing based on multi-scale visual transformer and attention mechanism

Jing Jin
Discover Artificial Intelligence
Human Pose and Action Recognition
article

Application of human action recognition in film and television editing based on multi-scale visual transformer and attention mechanism

Jing Jin
article en

Abstract

Abstract With the rapid growth of the scale of film and television content and the increasing demand for intelligent editing processes, problems such as large time scale spans, strong spatial interference, and weak continuity of action semantics in human actions in film and television works have become increasingly prominent. Therefore, the research aims to construct a human action recognition framework that integrates multi-scale visual Transformer and spatial–temporal attention mechanisms. Multi-scale context encoding is used to describe the long-term and short-term temporal correlation of actions, and multi-scale spatial and temporal attention networks are introduced to explicitly enhance key human body regions and key action fragments. Cross-scale feature fusion and action aggregation are combined to achieve robust discrimination. When the occlusion ratio reached 30%, the model recognition performance remained above 0.86, and the fluctuation amplitude was controlled within 0.02. In the transitional action detection task, the key action hit rate reached the highest 0.96, showing the ability to effectively capture the action switching process. The average detection delay was stable at around 30 ms, which was about 20% lower than that of the comparison method. The proposed multi-scale modeling and attention collaboration strategy can effectively improve the accuracy, stability and robustness of human action recognition in complex film and television scenes, providing a feasible technical path for intelligent film and television editing and content understanding.

Discover Artificial IntelligenceVol. 6(1)
Hunan University (CN)
Peace, Justice and strong institutions, Reduced inequalities
Openalex Percentile: Top 12%
Human Pose and Action Recognition
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Application of human action recognition in film and television editing based on multi-scale visual transformer and attention mechanism — Jing Jin · Discover Artificial Intelligence (2026) | TGRS Research Map | TGRS