Can large language models reproduce higher education grade bands? Cross-model study of calibration and grading bias in authentic student writing
Large Language Models (LLMs) are becoming increasingly used to support higher education assessment, yet evidence of their capability on reproducing authentic institutional grade-band labels remains limited. There is also limited understanding of how different model families vary in their grading behaviour, calibration, and tendency to produce systematic bias under consistent assessment conditions. This study examines whether LLMs can accurately grade university level student assignments using eight distinct pre-trained LLMs. The approach was tested using 114 human-marked university essays, and found that performance differed significantly across model configurations with some models aligning more closely with the human-grades while others showed clear patterns of undergrading. In some cases, models predicted Pass or Fail despite the essays being human-labelled as Merit or Distinction, showing the risk of systematic grading bias under consistent assessment conditions. Overall, the study indicates that large language models may be useful as supervised assessment-support tools, but they are not yet ready to be used independently for marking higher education student writing. Any future use in higher-stakes assessment would require further fine-tuning, a more directive assessment-specific knowledge base, and sustained human supervision.
Authors
- Aziza Mahomed (ORCID: https://orcid.org/0000-0002-6804-9383)
- Daniel L. Donaldson (ORCID: https://orcid.org/0000-0003-3419-3624)
- Logayna Kerwat (ORCID: https://orcid.org/0009-0000-7929-6390)
Institutions
- University College Birmingham (GB)
- University of Birmingham (GB)
Publication Details
- Journal
- Assessment & Evaluation in Higher Education
- Published
- 2026-09-21
- DOI
- https://doi.org/10.1080/02602938.2026.2734796
- Primary Topic
- Second Language Acquisition and Learning
- Type
- article
- Field-Weighted Citation Impact
- 0.00