Explainable Cross-Dataset Violence Detection for Public Safety Using a Two-Stream I3D Network

Camera networks now cover almost every public space, but the footage they generate has long outpaced what any human team can watch in real time. That gap is exactly where automated violence detection earns its place. The hard part is ambiguity—a hug or a bit of play-fighting can look a lot like a real attack, and systems that get this wrong flood operators with false alarms. Since violence lives in motion as much as in appearance, we build GoogLeNet-based Inflated 3D (I3D) two-stream classifiers that read space and time together. Three variants are trained and validated on the Real Life Violence Situations (RLVS) dataset (2000 videos, 80/20 split), each differing in temporal depth and resolution. The strongest, I3D Network 3, reaches 98.75% in-domain validation accuracy, with an F1 score of 98.74% and a Matthews Correlation Coefficient of 97.51%, values that should be read as upper bounds because of the file-level split and the validation-based checkpoint selection. Numbers on the home dataset only tell part of the story, so we push every model onto the independent RWF-2000 set, untouched during training. Even under this strict zero-shot transfer, I3D Network 3 holds 79.75% accuracy, although it is seen that this single value hides a strongly asymmetric error profile. The recall of the violence class decreases to 65.70%, in other words 343 of 1000 violent clips are missed, while the recall of the peaceful class remains at 93.80%, and McNemar and paired t-tests confirm the gaps between networks are real, rather than noise. For this reason, this result is evaluated as partial generalization. The only published values obtained under the same RLVS to RWF-2000 zero-shot protocol are 68.76% and 74.68%, so the margin is real but narrow. To open the black box, 3D Grad-CAM produces separate spatial and motion explanations that show where and when the model sees violence. Finally, the best model is embedded in ViolenceNet, a lightweight interface that keeps a human in the loop. Considering the missed-incident rate measured under domain shift, ViolenceNet is positioned as an assisted-review and triage tool, rather than as an autonomous alarm system.

Authors

Institutions

Publication Details

Journal
Applied Sciences
Published
2026-09-24
DOI
https://doi.org/10.3390/app16199524
Primary Topic
Human Pose and Action Recognition
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Explainable Cross-Dataset Violence Detection for Public Safety Using a Two-Stream I3D Network

Ufuk Bal, Kubilay Muhammed Sünnetci, Ahmet Çağdaş Seçkin
Applied Sciences
Human Pose and Action Recognition
article

Explainable Cross-Dataset Violence Detection for Public Safety Using a Two-Stream I3D Network

Ufuk Bal, Kubilay Muhammed Sünnetci, Ahmet Çağdaş Seçkin
article en

Abstract

Camera networks now cover almost every public space, but the footage they generate has long outpaced what any human team can watch in real time. That gap is exactly where automated violence detection earns its place. The hard part is ambiguity—a hug or a bit of play-fighting can look a lot like a real attack, and systems that get this wrong flood operators with false alarms. Since violence lives in motion as much as in appearance, we build GoogLeNet-based Inflated 3D (I3D) two-stream classifiers that read space and time together. Three variants are trained and validated on the Real Life Violence Situations (RLVS) dataset (2000 videos, 80/20 split), each differing in temporal depth and resolution. The strongest, I3D Network 3, reaches 98.75% in-domain validation accuracy, with an F1 score of 98.74% and a Matthews Correlation Coefficient of 97.51%, values that should be read as upper bounds because of the file-level split and the validation-based checkpoint selection. Numbers on the home dataset only tell part of the story, so we push every model onto the independent RWF-2000 set, untouched during training. Even under this strict zero-shot transfer, I3D Network 3 holds 79.75% accuracy, although it is seen that this single value hides a strongly asymmetric error profile. The recall of the violence class decreases to 65.70%, in other words 343 of 1000 violent clips are missed, while the recall of the peaceful class remains at 93.80%, and McNemar and paired t-tests confirm the gaps between networks are real, rather than noise. For this reason, this result is evaluated as partial generalization. The only published values obtained under the same RLVS to RWF-2000 zero-shot protocol are 68.76% and 74.68%, so the margin is real but narrow. To open the black box, 3D Grad-CAM produces separate spatial and motion explanations that show where and when the model sees violence. Finally, the best model is embedded in ViolenceNet, a lightweight interface that keeps a human in the loop. Considering the missed-incident rate measured under domain shift, ViolenceNet is positioned as an assisted-review and triage tool, rather than as an autonomous alarm system.

Applied SciencesVol. 16(19)
Osmaniye Korkut Ata University (TR), Adnan Menderes University (TR)
Peace, Justice and strong institutions
Openalex Percentile: Top 14%
Human Pose and Action Recognition
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.