Explainable Cross-Dataset Violence Detection for Public Safety Using a Two-Stream I3D Network
Camera networks now cover almost every public space, but the footage they generate has long outpaced what any human team can watch in real time. That gap is exactly where automated violence detection earns its place. The hard part is ambiguity—a hug or a bit of play-fighting can look a lot like a real attack, and systems that get this wrong flood operators with false alarms. Since violence lives in motion as much as in appearance, we build GoogLeNet-based Inflated 3D (I3D) two-stream classifiers that read space and time together. Three variants are trained and validated on the Real Life Violence Situations (RLVS) dataset (2000 videos, 80/20 split), each differing in temporal depth and resolution. The strongest, I3D Network 3, reaches 98.75% in-domain validation accuracy, with an F1 score of 98.74% and a Matthews Correlation Coefficient of 97.51%, values that should be read as upper bounds because of the file-level split and the validation-based checkpoint selection. Numbers on the home dataset only tell part of the story, so we push every model onto the independent RWF-2000 set, untouched during training. Even under this strict zero-shot transfer, I3D Network 3 holds 79.75% accuracy, although it is seen that this single value hides a strongly asymmetric error profile. The recall of the violence class decreases to 65.70%, in other words 343 of 1000 violent clips are missed, while the recall of the peaceful class remains at 93.80%, and McNemar and paired t-tests confirm the gaps between networks are real, rather than noise. For this reason, this result is evaluated as partial generalization. The only published values obtained under the same RLVS to RWF-2000 zero-shot protocol are 68.76% and 74.68%, so the margin is real but narrow. To open the black box, 3D Grad-CAM produces separate spatial and motion explanations that show where and when the model sees violence. Finally, the best model is embedded in ViolenceNet, a lightweight interface that keeps a human in the loop. Considering the missed-incident rate measured under domain shift, ViolenceNet is positioned as an assisted-review and triage tool, rather than as an autonomous alarm system.
Authors
- Ufuk Bal (ORCID: https://orcid.org/0000-0003-0345-6989)
- Kubilay Muhammed Sünnetci (ORCID: https://orcid.org/0000-0002-3500-5640)
- Ahmet Çağdaş Seçkin (ORCID: https://orcid.org/0000-0002-9849-3338)
Institutions
- Osmaniye Korkut Ata University (TR)
- Adnan Menderes University (TR)
Publication Details
- Journal
- Applied Sciences
- Published
- 2026-09-24
- DOI
- https://doi.org/10.3390/app16199524
- Primary Topic
- Human Pose and Action Recognition
- Type
- article
- Field-Weighted Citation Impact
- 0.00