Trained Still Wins: Narrowing the Gap to Zero-Shot Video Anomaly Detection
Video anomaly detection has two very different answers to the question of how a detector should acquire its notion of “normal” for a given camera: adapt its parameters to hours of normal footage recorded by that exact camera, or perform no target-scene adaptation at all and rely on frozen, heavily pretrained backbones—a vision-language image–text model and a large language model applied to text descriptions of frames. We call the second setting target-scene-training-free: the backbones themselves are trained on very large general-purpose corpora, but no parameter is updated, fine-tuned, or transfer-learned on the target scene. That setting is far more convenient to deploy, but how much accuracy does it give up, and can any of that gap be closed without target-scene adaptation? We define an anomaly operationally, as a frame whose score under a scoring function s(·)∈[0,1] exceeds a threshold, and evaluate every system with one metric definition and one implementation: frame-level area under the ROC curve (AUC) and equal error rate (EER), computed over all ground-truth-labeled test frames of each benchmark. We train a simplified future-frame-prediction network—a U-Net optimized with an intensity and gradient-difference loss, with the optical-flow and adversarial terms of the original design removed—on the normal-only training split of UCSD Ped1, UCSD Ped2, and CUHK Avenue, reaching a mean AUC of 0.848. A target-scene-training-free system starts far behind at 0.567, because its semantic signal is dropped in practice: the vision-language model that would produce it is too slow to run over a full test set. We close this gap in three ways, all evaluated on full test sets: CLIP-Guided Semantic Grounding (CSG), Vision-Language Reasoning (VLR), and Statistically-Calibrated Semantic Grounding (SCSG). Our best method raises mean AUC to 0.653, closing roughly a third of the gap with no target-scene adaptation. The trained model nonetheless leads on every benchmark, and because it is deliberately simplified, the gap we report is a conservative lower bound on the gap a fully engineered trained model would show. We report every result honestly, including a prompt-sensitivity analysis showing how much of the training-free result depends on prompt wording, cases where a more sophisticated method did not beat a simpler one, and a measured explanation of why UCSD Ped1 is the weakest benchmark for the trained model.
Authors
- Prasad B. Honnavalli (ORCID: https://orcid.org/0000-0001-7493-6221)
- Shylaja Ss (ORCID: https://orcid.org/0000-0003-2628-8973)
- Preet Kanwal (ORCID: https://orcid.org/0000-0002-7490-0090)
Institutions
- PES University (IN)
Publication Details
- Journal
- Electronics
- Published
- 2026-09-14
- DOI
- https://doi.org/10.3390/electronics15184172
- Primary Topic
- Anomaly Detection Techniques and Applications
- Type
- article
- Field-Weighted Citation Impact
- 0.00