Increasing the Efficiency of Comparative Judgment for Writing Assessment: Warm‐Starting Estimation and Selection Using NLP

Abstract Comparative judgment (CJ) is an assessment method in which assessors compare pairs of essays and judge which is of higher quality. While CJ produces valid and reliable results, it often requires many judgments before quality scores become reliable. In this study, we examine whether natural language processing (NLP) can improve the efficiency of CJ by warm‐starting the estimation of quality scores and the selection of essay pairs. We predict essay quality by fine‐tuning a transformer model and use these predictions to construct informative priors in a Bayesian Bradley‐Terry‐Luce model. We then introduce three selection rules based on predicted scores and essay embeddings. In this study, we focus on Dutch writing assessments in secondary and higher education in two settings: one‐off assessments and recurrent assessments. Warm‐start estimation reduced the number of judgments required to reach a reliability of .70 by approximately 40%‐68% compared with likelihood‐based cold‐start estimation and by approximately 24%‐64% compared with Bayesian cold‐start estimation. To reach a reliability of .90, warm‐start estimation yielded smaller reductions of around 15% compared with cold‐start estimation. The warm‐start selection rules we constructed yielded smaller additional gains. In practice, a reliability of .70 was reached after two to four judgments per essay, and a reliability of .90 after seven to fourteen judgments. Overall, warm‐starting CJ substantially reduces assessor workload and improves the practical applicability of CJ as a summative assessment method.

Authors

Institutions

Publication Details

Journal
Journal of Educational Measurement
Published
2026-09-30
DOI
https://doi.org/10.1111/jedm.70065
Primary Topic
Psychometric Methodologies and Testing
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Increasing the Efficiency of Comparative Judgment for Writing Assessment: Warm‐Starting Estimation and Selection Using NLP

Michiel De Vrindt, Renske Bouwer, Wim Van Den Noortgate, Anaïs Tack et al.
Journal of Educational Measurement
Psychometric Methodologies and Testing
article

Increasing the Efficiency of Comparative Judgment for Writing Assessment: Warm‐Starting Estimation and Selection Using NLP

Michiel De Vrindt, Renske Bouwer, Wim Van Den Noortgate, Anaïs Tack, Marije Lesterhuis
article en

Abstract

Abstract Comparative judgment (CJ) is an assessment method in which assessors compare pairs of essays and judge which is of higher quality. While CJ produces valid and reliable results, it often requires many judgments before quality scores become reliable. In this study, we examine whether natural language processing (NLP) can improve the efficiency of CJ by warm‐starting the estimation of quality scores and the selection of essay pairs. We predict essay quality by fine‐tuning a transformer model and use these predictions to construct informative priors in a Bayesian Bradley‐Terry‐Luce model. We then introduce three selection rules based on predicted scores and essay embeddings. In this study, we focus on Dutch writing assessments in secondary and higher education in two settings: one‐off assessments and recurrent assessments. Warm‐start estimation reduced the number of judgments required to reach a reliability of .70 by approximately 40%‐68% compared with likelihood‐based cold‐start estimation and by approximately 24%‐64% compared with Bayesian cold‐start estimation. To reach a reliability of .90, warm‐start estimation yielded smaller reductions of around 15% compared with cold‐start estimation. The warm‐start selection rules we constructed yielded smaller additional gains. In practice, a reliability of .70 was reached after two to four judgments per essay, and a reliability of .90 after seven to fourteen judgments. Overall, warm‐starting CJ substantially reduces assessor workload and improves the practical applicability of CJ as a summative assessment method.

Journal of Educational MeasurementVol. 63(4)
University of Applied Sciences Utrecht (NL), Utrecht University (NL), Imec the Netherlands (NL), IMEC (BE)
Quality Education
Openalex Percentile: Top 7%
Psychometric Methodologies and Testing
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.