Generative AI and LLMs in educational assessment: challenges, evidence, and a process-based framework

Introduction: This article examines how large language models (LLMs) can support educational assessment when verification, contextualisation, and human oversight are embedded throughout the process. Generative artificial intelligence (GenAI) is increasingly used in task design, grading, and feedback, but rapid model development has made some early findings partly obsolete. Materials and methods: The article uses an integrative synthesis design, combining the recent literature with the author’s previously published studies, new empirical material and additional analyses. Its original contribution is a three-stage sociotechnical framework: pre-assessment, covering authorised inputs, preprocessing, rubric alignment, model configuration, and input-fidelity checks; assessment, covering structured and repeated rubric-based scoring and discrepancy review; and post-assessment, covering personalised feedback and student perceptions. Results: The synthesis indicates that these controls can improve consistency, transparency, and pedagogical alignment of LLM-supported educational assessment while reducing, but not eliminating, hallucination, variability, bias, and over-reliance on model outputs. Evidence is strongest for supervised grading assistance, consistency checking, and formative feedback rather than autonomous high-stakes assessment. Conclusions: The three-stage sociotechnical GenAI-based assessment framework is adaptable across settings, but its performance remains dependent on the model, version, language, task, rubric, and institutional context. Recent developments in LLM architectures, including skill-based operations, the Model Context Protocol (MCP), secure connections to external data sources, enterprise systems and development tools, and wider context windows, suggest that the assessment of students’ written work can become more secure, systematic, and practical to implement.

Authors

Institutions

Publication Details

Journal
Academia AI and Applications
Published
2026-09-24
DOI
https://doi.org/10.20935/acadai8542
Primary Topic
Intelligent Tutoring Systems and Adaptive Learning
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Generative AI and LLMs in educational assessment: challenges, evidence, and a process-based framework

Jussi S. Jauhiainen
Academia AI and Applications
Intelligent Tutoring Systems and Adaptive Learning
article

Generative AI and LLMs in educational assessment: challenges, evidence, and a process-based framework

Jussi S. Jauhiainen
article en

Abstract

Introduction: This article examines how large language models (LLMs) can support educational assessment when verification, contextualisation, and human oversight are embedded throughout the process. Generative artificial intelligence (GenAI) is increasingly used in task design, grading, and feedback, but rapid model development has made some early findings partly obsolete. Materials and methods: The article uses an integrative synthesis design, combining the recent literature with the author’s previously published studies, new empirical material and additional analyses. Its original contribution is a three-stage sociotechnical framework: pre-assessment, covering authorised inputs, preprocessing, rubric alignment, model configuration, and input-fidelity checks; assessment, covering structured and repeated rubric-based scoring and discrepancy review; and post-assessment, covering personalised feedback and student perceptions. Results: The synthesis indicates that these controls can improve consistency, transparency, and pedagogical alignment of LLM-supported educational assessment while reducing, but not eliminating, hallucination, variability, bias, and over-reliance on model outputs. Evidence is strongest for supervised grading assistance, consistency checking, and formative feedback rather than autonomous high-stakes assessment. Conclusions: The three-stage sociotechnical GenAI-based assessment framework is adaptable across settings, but its performance remains dependent on the model, version, language, task, rubric, and institutional context. Recent developments in LLM architectures, including skill-based operations, the Model Context Protocol (MCP), secure connections to external data sources, enterprise systems and development tools, and wider context windows, suggest that the assessment of students’ written work can become more secure, systematic, and practical to implement.

Academia AI and ApplicationsVol. 2(3)
University of Turku (FI), University of Tartu (EE)
Quality Education
Openalex Percentile: Top 9%
Intelligent Tutoring Systems and Adaptive Learning
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Generative AI and LLMs in educational assessment: challenges, evidence, and a process-based framework — Jussi S. Jauhiainen · Academia AI and Applications (2026) | TGRS Research Map | TGRS