RMTA-Net: recurrent conditioned memory-guided temporal attention with spatial feature enhancement for video anomaly detection

Video anomaly detection aims to identify rare, unexpected events in real-world surveillance environments, where the diversity and scarcity of annotated examples make exhaustive supervision impractical. Prediction-based unsupervised learning addresses this problem by modeling normal spatiotemporal patterns and flagging anomalies through prediction deviations. However, existing approaches exhibit limited spatial representation capability and insufficient understanding of clip-level temporal dependency due to appearance variations and complex motion dynamics within the scene. To address these challenges, this paper introduces RMTA-Net, a future-frame prediction network for anomaly detection that jointly learns spatial representations and temporally coherent dependencies of sequential video frames. First, a residual spatial feature enhancement network progressively extracts structural and appearance information at multiple levels. The independently encoded spatial features are organized as an explicit temporal representation. Subsequently, a recurrent conditioned memory-guided temporal attention (RMTA) module integrates recurrent temporal processing with learnable memory banks and memory-guided temporal attention to model inter-frame dependencies within the observed clip. The dual-branch pipeline simultaneously processes the recurrent network output using attention-based memory access and temporal attention-driven global normality-prior retrieval. The learnable memory banks encode dataset-level normality priors, and adaptive gating fuses retrieved recurrent memory with contextual temporal relationships. Finally, an attention-enhanced decoder predicts future frames from the learned spatiotemporal embeddings, where anomalies are identified using prediction discrepancies. Extensive experiments were conducted on three benchmark datasets, namely UCSD Ped2, CUHK Avenue, and ShanghaiTech, demonstrating the effectiveness of the proposed method. RMTA-Net achieved frame-level AUC scores of 99.0%, 90.1%, and 75.8%, respectively, and remained competitive with several state-of-the-art methods.

Authors

Institutions

Publication Details

Journal
Complex & Intelligent Systems
Published
2026-10-05
DOI
https://doi.org/10.1007/s40747-026-02506-x
Primary Topic
Anomaly Detection Techniques and Applications
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

RMTA-Net: recurrent conditioned memory-guided temporal attention with spatial feature enhancement for video anomaly detection

Narinder Singh Punn, Mahua Bhattacharya, Santosh Prakash Chouhan
Complex & Intelligent Systems
Anomaly Detection Techniques and Applications
article

RMTA-Net: recurrent conditioned memory-guided temporal attention with spatial feature enhancement for video anomaly detection

Narinder Singh Punn, Mahua Bhattacharya, Santosh Prakash Chouhan
article en

Abstract

Video anomaly detection aims to identify rare, unexpected events in real-world surveillance environments, where the diversity and scarcity of annotated examples make exhaustive supervision impractical. Prediction-based unsupervised learning addresses this problem by modeling normal spatiotemporal patterns and flagging anomalies through prediction deviations. However, existing approaches exhibit limited spatial representation capability and insufficient understanding of clip-level temporal dependency due to appearance variations and complex motion dynamics within the scene. To address these challenges, this paper introduces RMTA-Net, a future-frame prediction network for anomaly detection that jointly learns spatial representations and temporally coherent dependencies of sequential video frames. First, a residual spatial feature enhancement network progressively extracts structural and appearance information at multiple levels. The independently encoded spatial features are organized as an explicit temporal representation. Subsequently, a recurrent conditioned memory-guided temporal attention (RMTA) module integrates recurrent temporal processing with learnable memory banks and memory-guided temporal attention to model inter-frame dependencies within the observed clip. The dual-branch pipeline simultaneously processes the recurrent network output using attention-based memory access and temporal attention-driven global normality-prior retrieval. The learnable memory banks encode dataset-level normality priors, and adaptive gating fuses retrieved recurrent memory with contextual temporal relationships. Finally, an attention-enhanced decoder predicts future frames from the learned spatiotemporal embeddings, where anomalies are identified using prediction discrepancies. Extensive experiments were conducted on three benchmark datasets, namely UCSD Ped2, CUHK Avenue, and ShanghaiTech, demonstrating the effectiveness of the proposed method. RMTA-Net achieved frame-level AUC scores of 99.0%, 90.1%, and 75.8%, respectively, and remained competitive with several state-of-the-art methods.

Complex & Intelligent Systems
Atal Bihari Vajpayee Indian Institute of Information Technology and Management (IN)
Openalex Percentile: Top 10%
Anomaly Detection Techniques and Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.