Hybrid deep learning for three-way classification of human-written, AI-generated, and AI-rephrased text
Abstract The rapid proliferation of large language models (LLMs) has made it increasingly difficult to distinguish AI-generated and AI-rephrased text from authentic human writing, posing serious risks to academic integrity, journalism, and e-commerce credibility. This paper presents a rigorous benchmark evaluation for three-way AI text origin classification designed to detect human-written, AI-generated, and AI-rephrased text by integrating handcrafted linguistic features, transformer-based contextual embeddings, and ensemble learning strategies. A corpus of 161,788 samples is curated from Amazon product reviews and Twitter posts, with AI-generated text produced via zero-shot prompting and AI-rephrased text produced via few-shot prompting using GPT, DeepSeek, and Kimi. The corpus comprises 46,844 human-written, 57,247 AI-generated, and 57,697 AI-rephrased samples distributed across training (113,249), validation (16,180), and test (32,359) splits. A hybrid feature vector of 168 dimensions is constructed, combining 68 handcrafted stylistic, lexical, linguistic, syntactic, and stylometric features with 100-dimensional embeddings derived from RoBERTa and the all-MiniLM-L6-v2 Sentence Transformer. Classical machine learning models, Random Forest, SVM, and Logistic Regression, are trained on this hybrid representation, while RoBERTa-base and DeBERTa-v3-base are fine-tuned end-to-end on raw tokenized text. Experimental evaluation demonstrates that transformer models substantially outperform classical methods, with RoBERTa achieving 95.49% accuracy and DeBERTa-v3 achieving 94.88%. A probability-averaging ensemble further improves accuracy to 95.66%, and a stacking ensemble with a Random Forest meta-learner reaches 96.46%. Ablation studies confirm the complementary value of handcrafted features alongside contextual embeddings. Across all models, AI-rephrased content is the most challenging category due to its semantic proximity to human writing, highlighting a persistent frontier for future research. The dataset and evaluation framework provide a reproducible foundation for future AI content detection research, with generalization to unseen generators and domains identified as primary directions for future work.
Authors
- Muhammad Shahzad Faisal (ORCID: https://orcid.org/0000-0003-1987-5782)
- Muhammad Saleem Khan (ORCID: https://orcid.org/0000-0003-2064-4690)
- Muhammad Ali Iqbal
- Najma Sadia
- Muhammad Naeem Khan
- Soo Kyun Kim
Institutions
- COMSATS University Islamabad (PK)
- Recep Tayyip Erdoğan University (TR)
- Jeju National University (KR)
Publication Details
- Journal
- Scientific Reports
- Published
- 2026-10-05
- DOI
- https://doi.org/10.1038/s41598-026-73841-9
- Primary Topic
- Authorship Attribution and Profiling
- Type
- article
- Field-Weighted Citation Impact
- 0.00
Funders
- Ministry of Education, India