The pitfalls of AI detection in academic writing: bias, false positives, and the need for inclusive assessment
Abstract AI-text detectors are increasingly used to police authorship in student assessment and scholarly publishing. This paper argues that they are the wrong instrument for that task. Its design combines a critical synthesis of empirical, information-theoretic, and policy evidence with an original documented exploratory multi-tool audit. The review synthesises empirical evaluations, information-theoretic results, and institutional policy decisions to develop a conceptual account of what detectors can directly observe: statistical or learned properties of submitted text, not provenance itself. Predictability-related properties such as low perplexity, reduced variation in local predictability, and formulaic density provide one mechanism through which both additional-language writing and highly conventional academic prose may become vulnerable to machine-authorship classification. While a growing number of universities across several national higher-education systems have disabled or restricted AI detection in high-stakes assessment, and a New York court has annulled a misconduct finding that rested on a detector score, detector reliance persists in three under-examined arenas: individual grading practice at scale, scholarly journal screening, and institutions without formal policy revision. The paper adds a documented exploratory multi-tool audit in which a single human-authored manuscript, submitted to five commercial detectors on the same day, received materially different classifications spanning from “0% human” to “Human Generated”, demonstrating substantial cross-tool inconsistency and limited independent verifiability of the underlying classifications. A paired exploratory comparison further found that extensive machine rewriting by a commercial “humanizer” materially degraded the manuscript while producing only a small change in the reported detection score. The paper distinguishes genuine false positives from definitional boundary cases involving AI-assisted translation and polishing, engages the counter-argument that AI-polished prose may itself flatten linguistic diversity in academic English, and proposes a concrete process-oriented assessment model based on declared starting points, versioned drafting, process artefacts, and targeted oral verification. The model predates generative AI in its component practices but acquires renewed relevance because final-text classification cannot establish authorship, understanding, or misconduct on its own.
Authors
- Victor Angelier
Institutions
- University of Essex (GB)
Publication Details
- Journal
- AI and Ethics
- Published
- 2026-09-18
- DOI
- https://doi.org/10.1007/s43681-026-01376-w
- Primary Topic
- Authorship Attribution and Profiling
- Type
- article
- Field-Weighted Citation Impact
- 0.00