Linda-Pro 1.3.1: a local ensemble detector of AI-generated text in English, Polish and Russian
Technical report on Linda-Pro 1.3.1: a local ensemble detector of AI-generated text in English, Polish and Russian (stylometry + two merged transformer models, one decision rule per language). Version 1.3.1 is a weights update retrained without text generated by Llama 3.x models, with additional permissively licensed human learner essays. On HumanizerBench (not used for training) 81.0% of humanized texts are flagged (version 1.1: 70%, 1.0: 31%); on a held-out October 2026 benchmark cycle 84.0%; on three unseen modern generators 95.5%. On held-out Polish and Russian sets the detector catches 79% and 80% of AI texts at 1.4% and 0.6% false positives. False positives on human texts are 0.2-0.4% on general, school and adult-learner English sets, 5.5% on TOEFL essays by non-native writers, and 2.0% on the MAGE human subset. Caveats and limitations (small Polish and Russian test sets, older small-model texts, monthly changing humanizers, not an independent audit) are reported in full. Version 1.1 report: https://doi.org/10.5281/zenodo.23080472 . Version 1.0 report: https://doi.org/10.5281/zenodo.23072494 . Code and evaluation script: https://github.com/Lendarixon/Linda . Self-published, not peer reviewed.
Authors
- Vladyslav Manzyuk
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-08
- DOI
- https://doi.org/10.5281/zenodo.23242014
- Primary Topic
- Authorship Attribution and Profiling
- Type
- article
- Field-Weighted Citation Impact
- 0.00