Comparative Analysis of Offensive Language in Political Discourse: A Machine Learning Approach Using Twitter Data

The proliferation of social media has transformed political communication, concurrently raising significant concerns regarding the prevalence of offensive and toxic language online. This study presents a comparative machine learning analysis of tweets posted by two prominent political figures: Geert Wilders (the Netherlands) and Boris Johnson (the United Kingdom). A dataset of 6,279 tweets was collected, preprocessed using a rigorous Natural Language Processing (NLP) pipeline, and vectorized using a Bag-of-Words (BoW) model augmented with a custom, domain-specific offensive lexicon. Seven machine learning classifiers were trained to perform authorship attribution, serving as a proxy for stylistic and lexical divergence. Empirical results indicate that the Artificial Neural Network (ANN) achieved the highest predictive performance with 96.82% accuracy, followed closely by Logistic Regression (96.74%). Conversely, K-Nearest Neighbors (KNN) exhibited poor performance (67.68%) due to the curse of dimensionality in high-dimensional text feature spaces. Furthermore, a lexicon-based frequency analysis revealed a higher prevalence of offensive content in the tweets of Geert Wilders compared to Boris Johnson. This paper details the methodology, model evaluation, and insights, while outlining future directions, including the integration of transformer-based models for deeper semantic understanding.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-12
DOI
https://doi.org/10.5281/zenodo.22720836
Primary Topic
Hate Speech and Cyberbullying Detection
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Comparative Analysis of Offensive Language in Political Discourse: A Machine Learning Approach Using Twitter Data

Sareena Bilal
Zenodo (CERN European Organization for Nuclear Research)
Hate Speech and Cyberbullying Detection
preprint

Comparative Analysis of Offensive Language in Political Discourse: A Machine Learning Approach Using Twitter Data

Sareena Bilal
preprint en

Abstract

The proliferation of social media has transformed political communication, concurrently raising significant concerns regarding the prevalence of offensive and toxic language online. This study presents a comparative machine learning analysis of tweets posted by two prominent political figures: Geert Wilders (the Netherlands) and Boris Johnson (the United Kingdom). A dataset of 6,279 tweets was collected, preprocessed using a rigorous Natural Language Processing (NLP) pipeline, and vectorized using a Bag-of-Words (BoW) model augmented with a custom, domain-specific offensive lexicon. Seven machine learning classifiers were trained to perform authorship attribution, serving as a proxy for stylistic and lexical divergence. Empirical results indicate that the Artificial Neural Network (ANN) achieved the highest predictive performance with 96.82% accuracy, followed closely by Logistic Regression (96.74%). Conversely, K-Nearest Neighbors (KNN) exhibited poor performance (67.68%) due to the curse of dimensionality in high-dimensional text feature spaces. Furthermore, a lexicon-based frequency analysis revealed a higher prevalence of offensive content in the tweets of Geert Wilders compared to Boris Johnson. This paper details the methodology, model evaluation, and insights, while outlining future directions, including the integration of transformer-based models for deeper semantic understanding.

Zenodo (CERN European Organization for Nuclear Research)
Quality Education
Hate Speech and Cyberbullying Detection
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Comparative Analysis of Offensive Language in Political Discourse: A Machine Learning Approach Using Twitter Data — Sareena Bilal · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS