Measurement Validity of AI-Text Detection in Turkish Academic Writing: Temporal, Provider, and Intervention Shifts

AI text detectors are used in academic integrity settings. Their validity may change across time, providers, and writing conditions. This study examined AI-text detection in Turkish academic writing using a leakage-audited historical corpus and a prespecified provider-known benchmark. The historical audit included 50,913 human-written documents, including 40,381 dated from 2000 to 2019. The benchmark contained 300 project-lineage-held-out human sources and 2,398 valid variants. Variants were produced by OpenAI, Gemini, DeepSeek, and Claude under four writing conditions: full generation, AI polishing, deeply mixed human–AI writing, and humanization. Six supervised detectors, one lineage-unresolved deployed system, and a frozen zero-shot baseline were evaluated at locked thresholds. Human false-positive rates and AI recall were estimated with exact binomial intervals. Score uncertainty was estimated with a source-seed clustered bootstrap. XLM-R achieved the highest held-out AUROC at 0.9154. Its human false-positive rate was 0.0400, with AI recall of 0.6952. BERTurk had a human false-positive rate of 0.0267 and AI recall of 0.6197. The deployed system reached 0.8603 recall and falsely flagged 35.0% of human texts. OpenAI outputs were the most difficult provider condition for all project-lineage-held-out supervised detectors. Deeply mixed text produced the lowest recall for the transformer models. AI-polished text was the most difficult condition for the TF-IDF baselines. The zero-shot baseline remained near chance at an AUROC of 0.5069. Its recall was zero at the locked threshold. Historical threshold transfer varied by architecture. The findings support local validation and explicit reporting of human false-positive rates in consequential academic use.

Authors

Institutions

Publication Details

Journal
Scientific journal of Mehmet Akif Ersoy University.
Published
2026-09-17
DOI
https://doi.org/10.70030/sjmakeu.2033004
Primary Topic
Academic integrity and plagiarism
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Measurement Validity of AI-Text Detection in Turkish Academic Writing: Temporal, Provider, and Intervention Shifts

Mahmut Sınecen, Mehmet Taha ESER, Mustafa Tüker
Scientific journal of Mehmet Akif Ersoy University.
Academic integrity and plagiarism
article

Measurement Validity of AI-Text Detection in Turkish Academic Writing: Temporal, Provider, and Intervention Shifts

Mahmut Sınecen, Mehmet Taha ESER, Mustafa Tüker
article en

Abstract

AI text detectors are used in academic integrity settings. Their validity may change across time, providers, and writing conditions. This study examined AI-text detection in Turkish academic writing using a leakage-audited historical corpus and a prespecified provider-known benchmark. The historical audit included 50,913 human-written documents, including 40,381 dated from 2000 to 2019. The benchmark contained 300 project-lineage-held-out human sources and 2,398 valid variants. Variants were produced by OpenAI, Gemini, DeepSeek, and Claude under four writing conditions: full generation, AI polishing, deeply mixed human–AI writing, and humanization. Six supervised detectors, one lineage-unresolved deployed system, and a frozen zero-shot baseline were evaluated at locked thresholds. Human false-positive rates and AI recall were estimated with exact binomial intervals. Score uncertainty was estimated with a source-seed clustered bootstrap. XLM-R achieved the highest held-out AUROC at 0.9154. Its human false-positive rate was 0.0400, with AI recall of 0.6952. BERTurk had a human false-positive rate of 0.0267 and AI recall of 0.6197. The deployed system reached 0.8603 recall and falsely flagged 35.0% of human texts. OpenAI outputs were the most difficult provider condition for all project-lineage-held-out supervised detectors. Deeply mixed text produced the lowest recall for the transformer models. AI-polished text was the most difficult condition for the TF-IDF baselines. The zero-shot baseline remained near chance at an AUROC of 0.5069. Its recall was zero at the locked threshold. Historical threshold transfer varied by architecture. The findings support local validation and explicit reporting of human false-positive rates in consequential academic use.

Scientific journal of Mehmet Akif Ersoy University.Vol. 9(1)
National Nuclear Research Center (AZ), Adnan Menderes University (TR)
Quality Education
Openalex Percentile: Top 7%
Academic integrity and plagiarism
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.