Telling Students to Evaluate Does Not Make It Happen: Task Stage, Offloading Tendency, and Error Detection in AI-Assisted Student Writing
Instructors can prescribe which part of a writing task a student may delegate to generative AI, but not how the student works. This four-arm randomised experiment separated them. At Site A, 744 undergraduates were randomised to one of four divisions of labour across four rounds of report writing—AI drafts and the student evaluates and revises; the student drafts and AI evaluates; the student drafts, AI revises, and the student adjudicates each change; no AI—with all AI withdrawn at Week 6 post-test, immediately after the final round (639 completed). A reduced three-arm replication at Site B randomised 126 students (109 completed). Both arrangements in which the student judged AI-produced text outperformed the arrangement in which AI judged student-produced text and the no-AI control (d = 0.35 to 0.36), whereas having AI evaluate was equivalent to no AI. The benefit was unevenly distributed (arm × tendency F(3, 629) = 7.25, p < 0.001), and logs showed students high in dependent tendency re-outsourcing their assigned evaluation (b = 0.072, p < 0.001). Satisfaction carried no information about error detection (r = 0.015), while self-rated quality was weakly associated with writing performance. Prescribing a division of labour governs the assignment, not the cognition.
Authors
- Mengmeng Fan (ORCID: https://orcid.org/0009-0005-6496-7426)
- Pengcheng Chang
Institutions
- Zhengzhou University (CN)
Publication Details
- Journal
- Behavioral Sciences
- Published
- 2026-09-17
- DOI
- https://doi.org/10.3390/bs16091671
- Primary Topic
- Artificial Intelligence in Healthcare and Education
- Type
- article
- Field-Weighted Citation Impact
- 0.00