Evaluating robustness and explainability in Arabic authorship models for detecting AI-generated texts

The rapid development of large language models has intensified the need for reliable methods that distinguish human-authored Arabic text from AI-generated text while remaining interpretable across orthographic, dialectal, and register variation. This study evaluates baseline, transformer, ensemble, and hybrid architectures for Arabic authorship attribution and AI-generated text detection using the AraGenEval corpus and a 10,000-text dialectal Arabic supplement. It combines in-domain evaluation, robustness testing, bootstrap confidence intervals, paired statistical comparisons, feature-based explainability, prediction-level error analysis, and probability-calibration diagnostics. The prediction-level matrix was used as the controlling source for exact confusion counts, class-level metrics, Brier scores, expected calibration error, and error-code distributions. The ensemble of independently fine-tuned AraBERT v2 and XLM-RoBERTa Large produced the strongest in-domain result among the tested systems (accuracy = 94.7%, weighted F1 = .947, macro F1 = .947, MCC = .894, 95% bootstrap CI [.932, .960]), but performance declined on diacritic-stripped text and sharply under dialectal/social-media shift. Exact calibration analysis showed that the Brier score worsened from .046 in-domain to .167 in the dialectal condition, while 10-bin expected calibration error increased to .163 for dialectal data. The hybrid AraBERT-stylometric model numerically exceeded AraBERT in the primary in-domain matrix (weighted F1 = .936 vs. .921), but the paired comparison was not statistically significant (p = .199); consequently, handcrafted stylometric features are interpreted as explanatory aids rather than evidence of reliable predictive superiority. The findings support calibrated, dialect-aware, and uncertainty-reported Arabic AI-text detection.

Authors

Institutions

Publication Details

Journal
Scientific Reports
Published
2026-09-11
DOI
https://doi.org/10.1038/s41598-026-67331-1
Primary Topic
Authorship Attribution and Profiling
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Evaluating robustness and explainability in Arabic authorship models for detecting AI-generated texts

Khaled Nasser Alfraidi, Bunder Sebail Alshammari, Mohammad Omar Aljudaiey, Walid Abdelhalim
Scientific Reports
Authorship Attribution and Profiling
article

Evaluating robustness and explainability in Arabic authorship models for detecting AI-generated texts

Khaled Nasser Alfraidi, Bunder Sebail Alshammari, Mohammad Omar Aljudaiey, Walid Abdelhalim
article en

Abstract

The rapid development of large language models has intensified the need for reliable methods that distinguish human-authored Arabic text from AI-generated text while remaining interpretable across orthographic, dialectal, and register variation. This study evaluates baseline, transformer, ensemble, and hybrid architectures for Arabic authorship attribution and AI-generated text detection using the AraGenEval corpus and a 10,000-text dialectal Arabic supplement. It combines in-domain evaluation, robustness testing, bootstrap confidence intervals, paired statistical comparisons, feature-based explainability, prediction-level error analysis, and probability-calibration diagnostics. The prediction-level matrix was used as the controlling source for exact confusion counts, class-level metrics, Brier scores, expected calibration error, and error-code distributions. The ensemble of independently fine-tuned AraBERT v2 and XLM-RoBERTa Large produced the strongest in-domain result among the tested systems (accuracy = 94.7%, weighted F1 = .947, macro F1 = .947, MCC = .894, 95% bootstrap CI [.932, .960]), but performance declined on diacritic-stripped text and sharply under dialectal/social-media shift. Exact calibration analysis showed that the Brier score worsened from .046 in-domain to .167 in the dialectal condition, while 10-bin expected calibration error increased to .163 for dialectal data. The hybrid AraBERT-stylometric model numerically exceeded AraBERT in the primary in-domain matrix (weighted F1 = .936 vs. .921), but the paired comparison was not statistically significant (p = .199); consequently, handcrafted stylometric features are interpreted as explanatory aids rather than evidence of reliable predictive superiority. The findings support calibrated, dialect-aware, and uncertainty-reported Arabic AI-text detection.

Scientific Reports
Beni-Suef University (EG), University of Ha'il (SA)
Quality Education
Openalex Percentile: Top 8%
Authorship Attribution and Profiling
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.