Comparative Analysis of Offensive Language in Political Discourse: A Machine Learning Approach Using Twitter Data
The proliferation of social media has transformed political communication, concurrently raising significant concerns regarding the prevalence of offensive and toxic language online. This study presents a comparative machine learning analysis of tweets posted by two prominent political figures: Geert Wilders (the Netherlands) and Boris Johnson (the United Kingdom). A dataset of 6,279 tweets was collected, preprocessed using a rigorous Natural Language Processing (NLP) pipeline, and vectorized using a Bag-of-Words (BoW) model augmented with a custom, domain-specific offensive lexicon. Seven machine learning classifiers were trained to perform authorship attribution, serving as a proxy for stylistic and lexical divergence. Empirical results indicate that the Artificial Neural Network (ANN) achieved the highest predictive performance with 96.82% accuracy, followed closely by Logistic Regression (96.74%). Conversely, K-Nearest Neighbors (KNN) exhibited poor performance (67.68%) due to the curse of dimensionality in high-dimensional text feature spaces. Furthermore, a lexicon-based frequency analysis revealed a higher prevalence of offensive content in the tweets of Geert Wilders compared to Boris Johnson. This paper details the methodology, model evaluation, and insights, while outlining future directions, including the integration of transformer-based models for deeper semantic understanding.
Authors
- Sareena Bilal
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-12
- DOI
- https://doi.org/10.5281/zenodo.22720836
- Primary Topic
- Hate Speech and Cyberbullying Detection
- Type
- preprint