An edge-adaptive vision-language framework for real-time open-set violence recognition in surveillance videos
Abstract Low-latency violent event detection is essential for intelligent edge surveillance. Current violent recognition models suffer three critical deployment drawbacks: prohibitive computational overhead from full fine‑tuning, inadequate temporal modeling for continuous video streams, and severe performance degradation when handling unseen violent categories in open surveillance scenes. We propose a lightweight open-set vision-language framework based on pre-trained CLIP. Its dual‑branch structure achieves cross‑modal alignment to support cross‑scene generalization toward novel violent patterns, and we further conduct preliminary explorations for open‑set violent identification for long monitoring footage. Evaluations on five surveillance datasets confirm competitive accuracy. Our parameter-efficient fine-tuning(PEFT) alleviates the inherent conflict between generalization and edge computing overhead, with strong adaptability across low-power embedded terminals. The framework establishes an extensible paradigm for real-time public safety early warning on edge devices, providing generalizable technical references for deploying multi-modal foundation models in urban security governance and broader edge vision perception tasks.
Authors
- Jiamu Bai
- Beibei Bian (ORCID: https://orcid.org/0009-0003-4935-5080)
- Donghui Zhang (ORCID: https://orcid.org/0000-0002-1690-4886)
Institutions
- Pennsylvania State University (US)
- Changchun Institute of Technology (CN)
Publication Details
- Journal
- Scientific Reports
- Published
- 2026-09-22
- DOI
- https://doi.org/10.1038/s41598-026-72763-w
- Primary Topic
- Human Pose and Action Recognition
- Type
- article
- Field-Weighted Citation Impact
- 0.00