English-Persian cross-language plagiarism detection: Multilingual deep learning approach
Cross-lingual plagiarism detection (CLPD) remains a critical challenge in Natural Language Processing, particularly for low-resource language pairs where translation and paraphrasing obscure source materials. This study introduces a robust framework for Persian-English CLPD by integrating state-of-the-art sentence representations with optimized machine learning classifiers. We utilized the BAAI/bge-m3 model to generate high-dimensional embeddings, employing cosine similarity to measure contrastive linguistic distances. To evaluate classification performance, we compared various architectures, including XGBoost, LSTM, and Logistic Regression (LR), across three categories: Exact, Semi-Exact, and Different. While deep learning and ensemble methods were initially considered, experimental results on an expanded dataset demonstrated that a fine-tuned LR model provided superior stability and generalization. The optimized LR approach significantly increased classification accuracy from 89% to 96% (p < 0.05). These findings suggest that high-quality embeddings can enhance detection performance in low-resource contexts without the computational overhead of complex architectures. To ensure transparency and reproducibility, our dataset and source code are available on GitHub.
Authors
- Marzieh Zarinbal (ORCID: https://orcid.org/0000-0001-9284-7329)
- Azadeh Mohebi (ORCID: https://orcid.org/0000-0002-6443-7450)
- Sara Ansar (ORCID: https://orcid.org/0009-0000-6226-7497)
Publication Details
- Journal
- PLoS ONE
- Published
- 2026-09-30
- DOI
- https://doi.org/10.1371/journal.pone.0354459
- Primary Topic
- Academic integrity and plagiarism
- Type
- article
- Field-Weighted Citation Impact
- 0.00