A multi-modal deep learning approach: towards robust sports activity recognition in aerial videos

Sports activity recognition in aerial videos has emerged to be a crucial area of study with the escalating use of uncrewed aerial vehicles (UAVs) in sports. Nevertheless, sports activity recognition in aerial views turns out to be an extremely difficult assignment because of small activity recognition targets, varying camera views, motion, occlusion, and multi-player sports activity recognition. Keeping these factors in mind, this study tends to introduce a multi-modal deep learning model for robust sports activity recognition in aerial videos. Based on this, the proposed model uses self-supervised learning visual features obtained via Self-DIstillation with NO labels version 2 (DINOv2), motion features obtained via dense optical flows, body geometry features via ellipsoidal modeling, pose features via multi-person pose estimation, and lastly, features related to social sports activity via graph modeling, and hypergraphs. These different features can be optimally combined via Bayesian optimization, with compact-bilinear pooling, which employs Graph SAmple and aggreGatE (GraphSAGE) to model the interpersonal associations, and ultimately, employs a gated recurrent unit (GRU) with a temporal attention model. This multi-modal deep learning model will be tested on two challenging athletic benchmarks: event recognition in aerial videos (ERA), which focuses on recognizing sports activity in the aerial context, and the other, named Collective Sports (C-Sports), which tends to recognize sports activity and group sports activity. Eventually, it becomes clear through the proposed model, which provides a top-1 accuracy measure of 59.8% (83.4% top-5), a top-1 measure of 93.8% on sports activity, and a top-1 measure of sports group activity recognition of 76.8%, which outperforms other methods and turns out to be close to other recent methods.

Authors

Institutions

Publication Details

Journal
PeerJ Computer Science
Published
2026-09-14
DOI
https://doi.org/10.7717/peerj-cs.4084
Primary Topic
Human Pose and Action Recognition
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

A multi-modal deep learning approach: towards robust sports activity recognition in aerial videos

Azzah Allahim, Bader Aldughayfiq, Nabil Almashfi, Dina Abdulaziz AlHammadi et al.
PeerJ Computer Science
Human Pose and Action Recognition
article

A multi-modal deep learning approach: towards robust sports activity recognition in aerial videos

Azzah Allahim, Bader Aldughayfiq, Nabil Almashfi, Dina Abdulaziz AlHammadi, Hisham Allahem, Ishrat Zahra, Ahmad Jalal
article en

Abstract

Sports activity recognition in aerial videos has emerged to be a crucial area of study with the escalating use of uncrewed aerial vehicles (UAVs) in sports. Nevertheless, sports activity recognition in aerial views turns out to be an extremely difficult assignment because of small activity recognition targets, varying camera views, motion, occlusion, and multi-player sports activity recognition. Keeping these factors in mind, this study tends to introduce a multi-modal deep learning model for robust sports activity recognition in aerial videos. Based on this, the proposed model uses self-supervised learning visual features obtained via Self-DIstillation with NO labels version 2 (DINOv2), motion features obtained via dense optical flows, body geometry features via ellipsoidal modeling, pose features via multi-person pose estimation, and lastly, features related to social sports activity via graph modeling, and hypergraphs. These different features can be optimally combined via Bayesian optimization, with compact-bilinear pooling, which employs Graph SAmple and aggreGatE (GraphSAGE) to model the interpersonal associations, and ultimately, employs a gated recurrent unit (GRU) with a temporal attention model. This multi-modal deep learning model will be tested on two challenging athletic benchmarks: event recognition in aerial videos (ERA), which focuses on recognizing sports activity in the aerial context, and the other, named Collective Sports (C-Sports), which tends to recognize sports activity and group sports activity. Eventually, it becomes clear through the proposed model, which provides a top-1 accuracy measure of 59.8% (83.4% top-5), a top-1 measure of 93.8% on sports activity, and a top-1 measure of sports group activity recognition of 76.8%, which outperforms other methods and turns out to be close to other recent methods.

PeerJ Computer ScienceVol. 12
Princess Nourah bint Abdulrahman University (SA), Korea University (KR), Jouf University (SA), Air University (PK)
Openalex Percentile: Top 13%
Human Pose and Action Recognition
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.