Action Recognition in Sports Videos Using Multiscale Convolutional Networks with Long Short-Term Memory (LSTM)-Based Temporal Modeling

Action recognition in sports videos remains challenging because of complex motion dynamics, occlusion, and high intra-class variability. Although existing deep learning approaches, including CNN-BiLSTM and transfer learning-based models, have demonstrated effectiveness in human activity recognition, their performance may be limited in sports scenarios with rapid, diverse movements. Many existing methods rely on single-scale convolutional filters, which may not effectively capture both fine-grained and coarse motion characteristics simultaneously. To address this limitation, this study proposes a Multi-Scale Convolutional Neural Network (MSCNN) integrated with a Long Short-Term Memory (LSTM) network for sports video action recognition. The MSCNN extracts spatial representations at multiple receptive fields through parallel convolutional kernels, enabling the learning of both detailed and contextual motion features. These features are subsequently processed by the LSTM to capture temporal dependencies and motion continuity across consecutive frames. Experimental evaluation was conducted on the UCF11, UCF Sports, and JHMDB benchmark datasets. The proposed MSCNN-LSTM model achieved classification accuracies of 98.3%, 95.4%, and 81.7%, respectively, outperforming the comparative approaches evaluated in this study. An ablation study further demonstrated the contribution of multi-scale feature extraction and temporal modeling to overall performance. These findings demonstrate the potential of the proposed framework to combine spatial and temporal information for sports video action recognition.

Authors

Institutions

Publication Details

Journal
Journal of Visualized Experiments
Published
2026-09-15
DOI
https://doi.org/10.3791/72030
Primary Topic
Human Pose and Action Recognition
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Action Recognition in Sports Videos Using Multiscale Convolutional Networks with Long Short-Term Memory (LSTM)-Based Temporal Modeling

Zhou Peng, Sarim Hafiz Mohd, Haiming Yang, Ma Xiaojuan
Journal of Visualized Experiments
Human Pose and Action Recognition
article

Action Recognition in Sports Videos Using Multiscale Convolutional Networks with Long Short-Term Memory (LSTM)-Based Temporal Modeling

Zhou Peng, Sarim Hafiz Mohd, Haiming Yang, Ma Xiaojuan
article en

Abstract

Action recognition in sports videos remains challenging because of complex motion dynamics, occlusion, and high intra-class variability. Although existing deep learning approaches, including CNN-BiLSTM and transfer learning-based models, have demonstrated effectiveness in human activity recognition, their performance may be limited in sports scenarios with rapid, diverse movements. Many existing methods rely on single-scale convolutional filters, which may not effectively capture both fine-grained and coarse motion characteristics simultaneously. To address this limitation, this study proposes a Multi-Scale Convolutional Neural Network (MSCNN) integrated with a Long Short-Term Memory (LSTM) network for sports video action recognition. The MSCNN extracts spatial representations at multiple receptive fields through parallel convolutional kernels, enabling the learning of both detailed and contextual motion features. These features are subsequently processed by the LSTM to capture temporal dependencies and motion continuity across consecutive frames. Experimental evaluation was conducted on the UCF11, UCF Sports, and JHMDB benchmark datasets. The proposed MSCNN-LSTM model achieved classification accuracies of 98.3%, 95.4%, and 81.7%, respectively, outperforming the comparative approaches evaluated in this study. An ablation study further demonstrated the contribution of multi-scale feature extraction and temporal modeling to overall performance. These findings demonstrate the potential of the proposed framework to combine spatial and temporal information for sports video action recognition.

Journal of Visualized Experiments(235)
Shandong University (CN), Shanghai Huayi Group (China) (CN), Dezhou University (CN), National University of Malaysia (MY)
Openalex Percentile: Top 42%
Human Pose and Action Recognition
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Action Recognition in Sports Videos Using Multiscale Convolutional Networks with Long Short-Term Memory (LSTM)-Based Temporal Modeling — Zhou Peng, Sarim Hafiz Mohd, et al. · Journal of Visualized Experiments (2026) | TGRS Research Map | TGRS