Multimodal deep learning framework for soccer video event detection and summarization

This paper presents a novel multimodal deep learning framework for soccer video analysis and summarization. The framework comprises four main modules: the Video Event Detection Module, the Keyword Spotting Module, the Crowd Cheering Classification Module, and the Video Summarization Module. The Video Event Detection Module integrates two submodules: the Shot Boundary Detection Submodule, which segments the match video into temporal shots, and the Video Classification Submodule, which classifies detected shots into soccer events. Trained on the SoccerVE dataset, this module achieves $$90.8\\%$$ accuracy in soccer event detection, outperforming the previous state of the art by 6.2 percentage points. The Keyword Spotting Module, trained on the SoccerAE dataset, detects match-related keywords from commentator audio and obtains $$85.8\\%$$ accuracy with fewer parameters than comparable models. The Crowd Cheering Classification Module, trained on the SoccerCC dataset, classifies spectator reactions and reaches $$98.7\\%$$ accuracy, effectively capturing match atmosphere. The Video Summarization Module integrates the outputs of the three event detection modules (Video Event Detection, Keyword Spotting, and Crowd Cheering Classification) and applies a scoring function with event weighting and knapsack optimization to generate concise and context-aware match summaries. Evaluation across ten complete matches demonstrates an overall summarization precision of $$91.0\\%$$ , representing a significant improvement over previous state-of-the-art methods. These results highlight the potential of multimodal deep learning to deliver high-precision and context-aware insights for sports analytics and automated highlight generation.

Authors

Institutions

Publication Details

Journal
Discover Artificial Intelligence
Published
2026-09-09
DOI
https://doi.org/10.1007/s44163-026-02093-3
Primary Topic
Video Analysis and Summarization
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Multimodal deep learning framework for soccer video event detection and summarization

M. Ghanbari, Ali Karimi
Discover Artificial Intelligence
Video Analysis and Summarization
article

Multimodal deep learning framework for soccer video event detection and summarization

M. Ghanbari, Ali Karimi
article en

Abstract

This paper presents a novel multimodal deep learning framework for soccer video analysis and summarization. The framework comprises four main modules: the Video Event Detection Module, the Keyword Spotting Module, the Crowd Cheering Classification Module, and the Video Summarization Module. The Video Event Detection Module integrates two submodules: the Shot Boundary Detection Submodule, which segments the match video into temporal shots, and the Video Classification Submodule, which classifies detected shots into soccer events. Trained on the SoccerVE dataset, this module achieves $$90.8\%$$ accuracy in soccer event detection, outperforming the previous state of the art by 6.2 percentage points. The Keyword Spotting Module, trained on the SoccerAE dataset, detects match-related keywords from commentator audio and obtains $$85.8\%$$ accuracy with fewer parameters than comparable models. The Crowd Cheering Classification Module, trained on the SoccerCC dataset, classifies spectator reactions and reaches $$98.7\%$$ accuracy, effectively capturing match atmosphere. The Video Summarization Module integrates the outputs of the three event detection modules (Video Event Detection, Keyword Spotting, and Crowd Cheering Classification) and applies a scoring function with event weighting and knapsack optimization to generate concise and context-aware match summaries. Evaluation across ten complete matches demonstrates an overall summarization precision of $$91.0\%$$ , representing a significant improvement over previous state-of-the-art methods. These results highlight the potential of multimodal deep learning to deliver high-precision and context-aware insights for sports analytics and automated highlight generation.

Discover Artificial IntelligenceVol. 6(1)
University of Essex (GB), University of Tehran (IR)
Openalex Percentile: Top 13%
Video Analysis and Summarization
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Multimodal deep learning framework for soccer video event detection and summarization — M. Ghanbari, Ali Karimi · Discover Artificial Intelligence (2026) | TGRS Research Map | TGRS