Trained Still Wins: Narrowing the Gap to Zero-Shot Video Anomaly Detection

Video anomaly detection has two very different answers to the question of how a detector should acquire its notion of “normal” for a given camera: adapt its parameters to hours of normal footage recorded by that exact camera, or perform no target-scene adaptation at all and rely on frozen, heavily pretrained backbones—a vision-language image–text model and a large language model applied to text descriptions of frames. We call the second setting target-scene-training-free: the backbones themselves are trained on very large general-purpose corpora, but no parameter is updated, fine-tuned, or transfer-learned on the target scene. That setting is far more convenient to deploy, but how much accuracy does it give up, and can any of that gap be closed without target-scene adaptation? We define an anomaly operationally, as a frame whose score under a scoring function s(·)∈[0,1] exceeds a threshold, and evaluate every system with one metric definition and one implementation: frame-level area under the ROC curve (AUC) and equal error rate (EER), computed over all ground-truth-labeled test frames of each benchmark. We train a simplified future-frame-prediction network—a U-Net optimized with an intensity and gradient-difference loss, with the optical-flow and adversarial terms of the original design removed—on the normal-only training split of UCSD Ped1, UCSD Ped2, and CUHK Avenue, reaching a mean AUC of 0.848. A target-scene-training-free system starts far behind at 0.567, because its semantic signal is dropped in practice: the vision-language model that would produce it is too slow to run over a full test set. We close this gap in three ways, all evaluated on full test sets: CLIP-Guided Semantic Grounding (CSG), Vision-Language Reasoning (VLR), and Statistically-Calibrated Semantic Grounding (SCSG). Our best method raises mean AUC to 0.653, closing roughly a third of the gap with no target-scene adaptation. The trained model nonetheless leads on every benchmark, and because it is deliberately simplified, the gap we report is a conservative lower bound on the gap a fully engineered trained model would show. We report every result honestly, including a prompt-sensitivity analysis showing how much of the training-free result depends on prompt wording, cases where a more sophisticated method did not beat a simpler one, and a measured explanation of why UCSD Ped1 is the weakest benchmark for the trained model.

Authors

Institutions

Publication Details

Journal
Electronics
Published
2026-09-14
DOI
https://doi.org/10.3390/electronics15184172
Primary Topic
Anomaly Detection Techniques and Applications
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Trained Still Wins: Narrowing the Gap to Zero-Shot Video Anomaly Detection

Prasad B. Honnavalli, Shylaja Ss, Preet Kanwal
Electronics
Anomaly Detection Techniques and Applications
article

Trained Still Wins: Narrowing the Gap to Zero-Shot Video Anomaly Detection

Prasad B. Honnavalli, Shylaja Ss, Preet Kanwal
article en

Abstract

Video anomaly detection has two very different answers to the question of how a detector should acquire its notion of “normal” for a given camera: adapt its parameters to hours of normal footage recorded by that exact camera, or perform no target-scene adaptation at all and rely on frozen, heavily pretrained backbones—a vision-language image–text model and a large language model applied to text descriptions of frames. We call the second setting target-scene-training-free: the backbones themselves are trained on very large general-purpose corpora, but no parameter is updated, fine-tuned, or transfer-learned on the target scene. That setting is far more convenient to deploy, but how much accuracy does it give up, and can any of that gap be closed without target-scene adaptation? We define an anomaly operationally, as a frame whose score under a scoring function s(·)∈[0,1] exceeds a threshold, and evaluate every system with one metric definition and one implementation: frame-level area under the ROC curve (AUC) and equal error rate (EER), computed over all ground-truth-labeled test frames of each benchmark. We train a simplified future-frame-prediction network—a U-Net optimized with an intensity and gradient-difference loss, with the optical-flow and adversarial terms of the original design removed—on the normal-only training split of UCSD Ped1, UCSD Ped2, and CUHK Avenue, reaching a mean AUC of 0.848. A target-scene-training-free system starts far behind at 0.567, because its semantic signal is dropped in practice: the vision-language model that would produce it is too slow to run over a full test set. We close this gap in three ways, all evaluated on full test sets: CLIP-Guided Semantic Grounding (CSG), Vision-Language Reasoning (VLR), and Statistically-Calibrated Semantic Grounding (SCSG). Our best method raises mean AUC to 0.653, closing roughly a third of the gap with no target-scene adaptation. The trained model nonetheless leads on every benchmark, and because it is deliberately simplified, the gap we report is a conservative lower bound on the gap a fully engineered trained model would show. We report every result honestly, including a prompt-sensitivity analysis showing how much of the training-free result depends on prompt wording, cases where a more sophisticated method did not beat a simpler one, and a measured explanation of why UCSD Ped1 is the weakest benchmark for the trained model.

ElectronicsVol. 15(18)
PES University (IN)
Openalex Percentile: Top 8%
Anomaly Detection Techniques and Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.