Inappropriate Content Classification Model for Digital Violence Detection Using Hybrid Data

This study presents an inappropriate-content classification framework for digital-violence detection in Ecuadorian Spanish, addressing extreme class imbalance and dialectal variation. Data were collected in a post-API setting through a Selenium-based scraping pipeline that reconstructs conversational context using a window of ? = 3 prior interventions. An initial zero-shot labeling attempt with LLMs (Hermes) revealed severe cultural misinterpretation, overestimating the Violence class by 50 times (97.5% false alerts), which motivated full human validation and targeted data engineering. To correct imbalance without contaminating evaluation, the corpus was split before augmentation (70/15/15), and minority classes were selectively leveled via few-shot generation with LLaMA 3.1, followed by strict deduplication and cosine-similarity filtering (? = 0.85) to preserve semantic diversity. Model selection compared BETO and mBERT, with BETO outperforming. Across four training scenarios, naïve oversampling produced artificially inflated metrics indicative of overfitting, whereas the proposed cost-sensitive and regularized configuration (BETO with semantic deduplication and weighted loss) achieved 94.39% accuracy, 0.9429 weighted F1, and 0.9022 macro F1, significantly improving recovery of critical classes. Results highlight that hybrid data are effective only when carefully curated and paired with leakage-free evaluation protocols. This work demonstrates how machine learning innovation and knowledge extraction from heterogeneous data can be combined to build robust models for digital-violence detection in low-resource, culturally specific contexts.

Authors

Institutions

Publication Details

Journal
Informatics
Published
2026-09-16
DOI
https://doi.org/10.3390/informatics13090151
Primary Topic
Hate Speech and Cyberbullying Detection
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Inappropriate Content Classification Model for Digital Violence Detection Using Hybrid Data

Carlos E. Anchundia, Patricio Zambrano, Marco Sánchez, Adrian Esteban Paguay Montenegro et al.
Informatics
Hate Speech and Cyberbullying Detection
article

Inappropriate Content Classification Model for Digital Violence Detection Using Hybrid Data

Carlos E. Anchundia, Patricio Zambrano, Marco Sánchez, Adrian Esteban Paguay Montenegro, Andrea Damarys Oña Calahorrano, Juan Sebastián León Espinosa, Johan Sebastian Illicachi Manzano
article en

Abstract

This study presents an inappropriate-content classification framework for digital-violence detection in Ecuadorian Spanish, addressing extreme class imbalance and dialectal variation. Data were collected in a post-API setting through a Selenium-based scraping pipeline that reconstructs conversational context using a window of ? = 3 prior interventions. An initial zero-shot labeling attempt with LLMs (Hermes) revealed severe cultural misinterpretation, overestimating the Violence class by 50 times (97.5% false alerts), which motivated full human validation and targeted data engineering. To correct imbalance without contaminating evaluation, the corpus was split before augmentation (70/15/15), and minority classes were selectively leveled via few-shot generation with LLaMA 3.1, followed by strict deduplication and cosine-similarity filtering (? = 0.85) to preserve semantic diversity. Model selection compared BETO and mBERT, with BETO outperforming. Across four training scenarios, naïve oversampling produced artificially inflated metrics indicative of overfitting, whereas the proposed cost-sensitive and regularized configuration (BETO with semantic deduplication and weighted loss) achieved 94.39% accuracy, 0.9429 weighted F1, and 0.9022 macro F1, significantly improving recovery of critical classes. Results highlight that hybrid data are effective only when carefully curated and paired with leakage-free evaluation protocols. This work demonstrates how machine learning innovation and knowledge extraction from heterogeneous data can be combined to build robust models for digital-violence detection in low-resource, culturally specific contexts.

InformaticsVol. 13(9)
National Polytechnic School (EC)
Openalex Percentile: Top 8%
Hate Speech and Cyberbullying Detection
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.