An edge-adaptive vision-language framework for real-time open-set violence recognition in surveillance videos

Abstract Low-latency violent event detection is essential for intelligent edge surveillance. Current violent recognition models suffer three critical deployment drawbacks: prohibitive computational overhead from full fine‑tuning, inadequate temporal modeling for continuous video streams, and severe performance degradation when handling unseen violent categories in open surveillance scenes. We propose a lightweight open-set vision-language framework based on pre-trained CLIP. Its dual‑branch structure achieves cross‑modal alignment to support cross‑scene generalization toward novel violent patterns, and we further conduct preliminary explorations for open‑set violent identification for long monitoring footage. Evaluations on five surveillance datasets confirm competitive accuracy. Our parameter-efficient fine-tuning(PEFT) alleviates the inherent conflict between generalization and edge computing overhead, with strong adaptability across low-power embedded terminals. The framework establishes an extensible paradigm for real-time public safety early warning on edge devices, providing generalizable technical references for deploying multi-modal foundation models in urban security governance and broader edge vision perception tasks.

Authors

Institutions

Publication Details

Journal
Scientific Reports
Published
2026-09-22
DOI
https://doi.org/10.1038/s41598-026-72763-w
Primary Topic
Human Pose and Action Recognition
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

An edge-adaptive vision-language framework for real-time open-set violence recognition in surveillance videos

Jiamu Bai, Beibei Bian, Donghui Zhang
Scientific Reports
Human Pose and Action Recognition
article

An edge-adaptive vision-language framework for real-time open-set violence recognition in surveillance videos

Jiamu Bai, Beibei Bian, Donghui Zhang
article en

Abstract

Abstract Low-latency violent event detection is essential for intelligent edge surveillance. Current violent recognition models suffer three critical deployment drawbacks: prohibitive computational overhead from full fine‑tuning, inadequate temporal modeling for continuous video streams, and severe performance degradation when handling unseen violent categories in open surveillance scenes. We propose a lightweight open-set vision-language framework based on pre-trained CLIP. Its dual‑branch structure achieves cross‑modal alignment to support cross‑scene generalization toward novel violent patterns, and we further conduct preliminary explorations for open‑set violent identification for long monitoring footage. Evaluations on five surveillance datasets confirm competitive accuracy. Our parameter-efficient fine-tuning(PEFT) alleviates the inherent conflict between generalization and edge computing overhead, with strong adaptability across low-power embedded terminals. The framework establishes an extensible paradigm for real-time public safety early warning on edge devices, providing generalizable technical references for deploying multi-modal foundation models in urban security governance and broader edge vision perception tasks.

Scientific Reports
Pennsylvania State University (US), Changchun Institute of Technology (CN)
Peace, Justice and strong institutions
Openalex Percentile: Top 14%
Human Pose and Action Recognition
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

An edge-adaptive vision-language framework for real-time open-set violence recognition in surveillance videos — Jiamu Bai, Beibei Bian, et al. · Scientific Reports (2026) | TGRS Research Map | TGRS