Can large language models reproduce higher education grade bands? Cross-model study of calibration and grading bias in authentic student writing

Large Language Models (LLMs) are becoming increasingly used to support higher education assessment, yet evidence of their capability on reproducing authentic institutional grade-band labels remains limited. There is also limited understanding of how different model families vary in their grading behaviour, calibration, and tendency to produce systematic bias under consistent assessment conditions. This study examines whether LLMs can accurately grade university level student assignments using eight distinct pre-trained LLMs. The approach was tested using 114 human-marked university essays, and found that performance differed significantly across model configurations with some models aligning more closely with the human-grades while others showed clear patterns of undergrading. In some cases, models predicted Pass or Fail despite the essays being human-labelled as Merit or Distinction, showing the risk of systematic grading bias under consistent assessment conditions. Overall, the study indicates that large language models may be useful as supervised assessment-support tools, but they are not yet ready to be used independently for marking higher education student writing. Any future use in higher-stakes assessment would require further fine-tuning, a more directive assessment-specific knowledge base, and sustained human supervision.

Authors

Institutions

Publication Details

Journal
Assessment & Evaluation in Higher Education
Published
2026-09-21
DOI
https://doi.org/10.1080/02602938.2026.2734796
Primary Topic
Second Language Acquisition and Learning
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Can large language models reproduce higher education grade bands? Cross-model study of calibration and grading bias in authentic student writing

Aziza Mahomed, Daniel L. Donaldson, Logayna Kerwat
Assessment & Evaluation in Higher Education
Second Language Acquisition and Learning
article

Can large language models reproduce higher education grade bands? Cross-model study of calibration and grading bias in authentic student writing

Aziza Mahomed, Daniel L. Donaldson, Logayna Kerwat
article en

Abstract

Large Language Models (LLMs) are becoming increasingly used to support higher education assessment, yet evidence of their capability on reproducing authentic institutional grade-band labels remains limited. There is also limited understanding of how different model families vary in their grading behaviour, calibration, and tendency to produce systematic bias under consistent assessment conditions. This study examines whether LLMs can accurately grade university level student assignments using eight distinct pre-trained LLMs. The approach was tested using 114 human-marked university essays, and found that performance differed significantly across model configurations with some models aligning more closely with the human-grades while others showed clear patterns of undergrading. In some cases, models predicted Pass or Fail despite the essays being human-labelled as Merit or Distinction, showing the risk of systematic grading bias under consistent assessment conditions. Overall, the study indicates that large language models may be useful as supervised assessment-support tools, but they are not yet ready to be used independently for marking higher education student writing. Any future use in higher-stakes assessment would require further fine-tuning, a more directive assessment-specific knowledge base, and sustained human supervision.

Assessment & Evaluation in Higher Education
University College Birmingham (GB), University of Birmingham (GB)
Quality Education
Openalex Percentile: Top 5%
Second Language Acquisition and Learning
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Can large language models reproduce higher education grade bands? Cross-model study of calibration and grading bias in authentic student writing — Aziza Mahomed, Daniel L. Donaldson, et al. · Assessment & Evaluation in Higher Education (2026) | TGRS Research Map | TGRS