English-Persian cross-language plagiarism detection: Multilingual deep learning approach

Cross-lingual plagiarism detection (CLPD) remains a critical challenge in Natural Language Processing, particularly for low-resource language pairs where translation and paraphrasing obscure source materials. This study introduces a robust framework for Persian-English CLPD by integrating state-of-the-art sentence representations with optimized machine learning classifiers. We utilized the BAAI/bge-m3 model to generate high-dimensional embeddings, employing cosine similarity to measure contrastive linguistic distances. To evaluate classification performance, we compared various architectures, including XGBoost, LSTM, and Logistic Regression (LR), across three categories: Exact, Semi-Exact, and Different. While deep learning and ensemble methods were initially considered, experimental results on an expanded dataset demonstrated that a fine-tuned LR model provided superior stability and generalization. The optimized LR approach significantly increased classification accuracy from 89% to 96% (p < 0.05). These findings suggest that high-quality embeddings can enhance detection performance in low-resource contexts without the computational overhead of complex architectures. To ensure transparency and reproducibility, our dataset and source code are available on GitHub.

Authors

Publication Details

Journal
PLoS ONE
Published
2026-09-30
DOI
https://doi.org/10.1371/journal.pone.0354459
Primary Topic
Academic integrity and plagiarism
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

English-Persian cross-language plagiarism detection: Multilingual deep learning approach

Marzieh Zarinbal, Azadeh Mohebi, Sara Ansar
PLoS ONE
Academic integrity and plagiarism
article

English-Persian cross-language plagiarism detection: Multilingual deep learning approach

Marzieh Zarinbal, Azadeh Mohebi, Sara Ansar
article en

Abstract

Cross-lingual plagiarism detection (CLPD) remains a critical challenge in Natural Language Processing, particularly for low-resource language pairs where translation and paraphrasing obscure source materials. This study introduces a robust framework for Persian-English CLPD by integrating state-of-the-art sentence representations with optimized machine learning classifiers. We utilized the BAAI/bge-m3 model to generate high-dimensional embeddings, employing cosine similarity to measure contrastive linguistic distances. To evaluate classification performance, we compared various architectures, including XGBoost, LSTM, and Logistic Regression (LR), across three categories: Exact, Semi-Exact, and Different. While deep learning and ensemble methods were initially considered, experimental results on an expanded dataset demonstrated that a fine-tuned LR model provided superior stability and generalization. The optimized LR approach significantly increased classification accuracy from 89% to 96% (p < 0.05). These findings suggest that high-quality embeddings can enhance detection performance in low-resource contexts without the computational overhead of complex architectures. To ensure transparency and reproducibility, our dataset and source code are available on GitHub.

PLoS ONEVol. 21(9)
Quality Education
Openalex Percentile: Top 7%
Academic integrity and plagiarism
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

English-Persian cross-language plagiarism detection: Multilingual deep learning approach — Marzieh Zarinbal, Azadeh Mohebi, et al. · PLoS ONE (2026) | TGRS Research Map | TGRS