Linda-Pro 1.3.1: a local ensemble detector of AI-generated text in English, Polish and Russian

Technical report on Linda-Pro 1.3.1: a local ensemble detector of AI-generated text in English, Polish and Russian (stylometry + two merged transformer models, one decision rule per language). Version 1.3.1 is a weights update retrained without text generated by Llama 3.x models, with additional permissively licensed human learner essays. On HumanizerBench (not used for training) 81.0% of humanized texts are flagged (version 1.1: 70%, 1.0: 31%); on a held-out October 2026 benchmark cycle 84.0%; on three unseen modern generators 95.5%. On held-out Polish and Russian sets the detector catches 79% and 80% of AI texts at 1.4% and 0.6% false positives. False positives on human texts are 0.2-0.4% on general, school and adult-learner English sets, 5.5% on TOEFL essays by non-native writers, and 2.0% on the MAGE human subset. Caveats and limitations (small Polish and Russian test sets, older small-model texts, monthly changing humanizers, not an independent audit) are reported in full. Version 1.1 report: https://doi.org/10.5281/zenodo.23080472 . Version 1.0 report: https://doi.org/10.5281/zenodo.23072494 . Code and evaluation script: https://github.com/Lendarixon/Linda . Self-published, not peer reviewed.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-08
DOI
https://doi.org/10.5281/zenodo.23242014
Primary Topic
Authorship Attribution and Profiling
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Linda-Pro 1.3.1: a local ensemble detector of AI-generated text in English, Polish and Russian

Vladyslav Manzyuk
Zenodo (CERN European Organization for Nuclear Research)
Authorship Attribution and Profiling
article

Linda-Pro 1.3.1: a local ensemble detector of AI-generated text in English, Polish and Russian

Vladyslav Manzyuk
article en

Abstract

Technical report on Linda-Pro 1.3.1: a local ensemble detector of AI-generated text in English, Polish and Russian (stylometry + two merged transformer models, one decision rule per language). Version 1.3.1 is a weights update retrained without text generated by Llama 3.x models, with additional permissively licensed human learner essays. On HumanizerBench (not used for training) 81.0% of humanized texts are flagged (version 1.1: 70%, 1.0: 31%); on a held-out October 2026 benchmark cycle 84.0%; on three unseen modern generators 95.5%. On held-out Polish and Russian sets the detector catches 79% and 80% of AI texts at 1.4% and 0.6% false positives. False positives on human texts are 0.2-0.4% on general, school and adult-learner English sets, 5.5% on TOEFL essays by non-native writers, and 2.0% on the MAGE human subset. Caveats and limitations (small Polish and Russian test sets, older small-model texts, monthly changing humanizers, not an independent audit) are reported in full. Version 1.1 report: https://doi.org/10.5281/zenodo.23080472 . Version 1.0 report: https://doi.org/10.5281/zenodo.23072494 . Code and evaluation script: https://github.com/Lendarixon/Linda . Self-published, not peer reviewed.

Zenodo (CERN European Organization for Nuclear Research)
Openalex Percentile: Top 12%
Authorship Attribution and Profiling
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.