Mapping offensive discourse in sports on X platform: a weakly supervised NLP and journalism-based analysis

Abstract Purpose Social media platforms have become a central place for sport-related public discourse that often involves emotionally charged and offensive language. This study examines offensive sports discourse on X (formerly Twitter) through an interdisciplinary approach that combines natural language processing (NLP) and journalism research. Design/methodology/approach A dataset of 7,279 sport-related tweets that were collected between January and December 2025 was analyzed. An initial set of 350 tweets were manually annotated first and then used to train a machine learning classifier based on TF-IDF features and logistic regression. The classifier got an accuracy of 0.71. A confidence-based weakly supervised labelling strategy was then applied, resulting in 2,042 high-confidence labelled tweets. Findings Quantitative analysis shows that offensive tweets demonstrated higher average engagement levels than regular tweets in terms of likes, retweets, views, and replies, although these differences were not statistically significant. Unverified users predominantly produce offensive content and this contents appear more frequently in original tweets than in replies. Lexical analysis reveals that offensive discourse combines explicit profanity with sports-related terminology. The findings demonstrate how NLP-based methods can support journalism research by revealing large scale patterns in digital sports communication, and how these methods can be used in interdisciplinary settings to provide unique and novel insights through the analysis of the available data. Practical implications This research assists the growing field of computational journalism by leveraging NLP techniques with journalism-based research in an interdisciplinary settings to gain unique and novel insights about the domain-specific online sports discourse. Social implications This study helps us to achieve empirical insights and intuitions into fan behavior, interaction and the pattern of their engagement. At the same time, it is very important to perform journalistic interpretation for the purpose of contextualizing computational findings within ethical, cultural, and professional frameworks. Originality/value This study adopts established NLP based methods to explore large scale communicative patterns in online sports discourse.

Authors

Institutions

Publication Details

Journal
Online Media and Global Communication
Published
2026-09-16
DOI
https://doi.org/10.1515/omgc-2026-0027
Primary Topic
Hate Speech and Cyberbullying Detection
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Mapping offensive discourse in sports on X platform: a weakly supervised NLP and journalism-based analysis

Ahmed A. O. Shbair, Tuba LİVBERBER, Melih Günay, Lutf Ul Rahman Haqmal
Online Media and Global Communication
Hate Speech and Cyberbullying Detection
article

Mapping offensive discourse in sports on X platform: a weakly supervised NLP and journalism-based analysis

Ahmed A. O. Shbair, Tuba LİVBERBER, Melih Günay, Lutf Ul Rahman Haqmal
article en

Abstract

Abstract Purpose Social media platforms have become a central place for sport-related public discourse that often involves emotionally charged and offensive language. This study examines offensive sports discourse on X (formerly Twitter) through an interdisciplinary approach that combines natural language processing (NLP) and journalism research. Design/methodology/approach A dataset of 7,279 sport-related tweets that were collected between January and December 2025 was analyzed. An initial set of 350 tweets were manually annotated first and then used to train a machine learning classifier based on TF-IDF features and logistic regression. The classifier got an accuracy of 0.71. A confidence-based weakly supervised labelling strategy was then applied, resulting in 2,042 high-confidence labelled tweets. Findings Quantitative analysis shows that offensive tweets demonstrated higher average engagement levels than regular tweets in terms of likes, retweets, views, and replies, although these differences were not statistically significant. Unverified users predominantly produce offensive content and this contents appear more frequently in original tweets than in replies. Lexical analysis reveals that offensive discourse combines explicit profanity with sports-related terminology. The findings demonstrate how NLP-based methods can support journalism research by revealing large scale patterns in digital sports communication, and how these methods can be used in interdisciplinary settings to provide unique and novel insights through the analysis of the available data. Practical implications This research assists the growing field of computational journalism by leveraging NLP techniques with journalism-based research in an interdisciplinary settings to gain unique and novel insights about the domain-specific online sports discourse. Social implications This study helps us to achieve empirical insights and intuitions into fan behavior, interaction and the pattern of their engagement. At the same time, it is very important to perform journalistic interpretation for the purpose of contextualizing computational findings within ethical, cultural, and professional frameworks. Originality/value This study adopts established NLP based methods to explore large scale communicative patterns in online sports discourse.

Online Media and Global Communication
Akdeniz University (TR)
Quality Education
Openalex Percentile: Top 8%
Hate Speech and Cyberbullying Detection
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.