A Transformer-based Hybrid Model for Context-aware Multilingual Comment Grading

Language processing across diverse linguistic systems presents significant challenges, especially when dealing with morphologically rich languages like Tamil and syntactically complex languages like German. Conventional methods struggle with code-mixed and multilingual inputs, and most existing comment analysis systems are confined to either offensive detection or single-language essay grading. However, existing multilingual comment analysis methods rarely provide a unified framework capable of jointly evaluating linguistic quality and contextual offensiveness in multilingual and code-mixed social media comments. These limitations highlight the absence of a robust framework capable of handling multilingual, context-driven comment quality evaluation. Addressing this gap, the present study introduces a dual-dimensional grading framework designed to jointly evaluate offensiveness and linguistic quality in social media comments written in Tamil, German, and English. The proposed architecture integrates a transformer-based multilingual encoder with an attention-driven classification layer, allowing the model to understand nuanced semantic and grammatical patterns specific to each language. A combination of XLM-RoBERTa for contextual embedding and BiLSTM with self-attention for classification was employed, enabling simultaneous learning of language-specific features and generalizable grading patterns. The model was trained and evaluated using real-world YouTube comment datasets in both Tamil and German, manually labeled across multiple quality and offensiveness levels. Extensive experiments were conducted to assess grading accuracy, semantic relevance, and model robustness, utilizing precision, recall, F1-score, and quadratic weighted kappa as evaluation metrics. The results demonstrated consistent performance gains, achieving higher grading precision and better agreement scores than traditional methods. The proposed model achieved 92.87% accuracy and 92.50% F1-score on the Tamil dataset, while attaining 94.21% accuracy and 94.04% F1-score on the German dataset. Comparative evaluations demonstrated that the proposed XLM-RoBERTa with BiLSTM architecture more effectively captures multilingual contextual semantics, transliterated expressions, and sequential language dependencies than conventional recurrent and transformer-based grading approaches.

Authors

Institutions

Publication Details

Journal
International Journal of Computational Intelligence Systems
Published
2026-10-04
DOI
https://doi.org/10.1007/s44196-026-01606-3
Primary Topic
Hate Speech and Cyberbullying Detection
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

A Transformer-based Hybrid Model for Context-aware Multilingual Comment Grading

S. Thivaharan, G. Srivatsun
International Journal of Computational Intelligence Systems
Hate Speech and Cyberbullying Detection
article

A Transformer-based Hybrid Model for Context-aware Multilingual Comment Grading

S. Thivaharan, G. Srivatsun
article en

Abstract

Language processing across diverse linguistic systems presents significant challenges, especially when dealing with morphologically rich languages like Tamil and syntactically complex languages like German. Conventional methods struggle with code-mixed and multilingual inputs, and most existing comment analysis systems are confined to either offensive detection or single-language essay grading. However, existing multilingual comment analysis methods rarely provide a unified framework capable of jointly evaluating linguistic quality and contextual offensiveness in multilingual and code-mixed social media comments. These limitations highlight the absence of a robust framework capable of handling multilingual, context-driven comment quality evaluation. Addressing this gap, the present study introduces a dual-dimensional grading framework designed to jointly evaluate offensiveness and linguistic quality in social media comments written in Tamil, German, and English. The proposed architecture integrates a transformer-based multilingual encoder with an attention-driven classification layer, allowing the model to understand nuanced semantic and grammatical patterns specific to each language. A combination of XLM-RoBERTa for contextual embedding and BiLSTM with self-attention for classification was employed, enabling simultaneous learning of language-specific features and generalizable grading patterns. The model was trained and evaluated using real-world YouTube comment datasets in both Tamil and German, manually labeled across multiple quality and offensiveness levels. Extensive experiments were conducted to assess grading accuracy, semantic relevance, and model robustness, utilizing precision, recall, F1-score, and quadratic weighted kappa as evaluation metrics. The results demonstrated consistent performance gains, achieving higher grading precision and better agreement scores than traditional methods. The proposed model achieved 92.87% accuracy and 92.50% F1-score on the Tamil dataset, while attaining 94.21% accuracy and 94.04% F1-score on the German dataset. Comparative evaluations demonstrated that the proposed XLM-RoBERTa with BiLSTM architecture more effectively captures multilingual contextual semantics, transliterated expressions, and sequential language dependencies than conventional recurrent and transformer-based grading approaches.

International Journal of Computational Intelligence Systems
PSG INSTITUTE OF TECHNOLOGY AND APPLIED RESEARCH (IN), KPR Institute of Engineering and Technology (IN)
Openalex Percentile: Top 11%
Hate Speech and Cyberbullying Detection
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.