Evaluating Theme-Specific Feedback Emphasis and Ratings in Human and Large Language Model Raters

Large Language Models (LLMs) have shown promise in automated writing evaluation that provides estimates of writing proficiency ratings and feedback, yet validation studies concentrate on score agreement, overlooking whether human and LLM raters attend to the same aspects in rubrics. We evaluate comparability between a human reference benchmark and an LLM rater using both agreement indices and Many-Facet Rasch Measurement (MFRM) in ratings and feedback, and propose the psychometric approach that separates what theme the rater emphasizes in feedback and how strongly the rater evaluates within a theme. We use an argumentative writing assessment with 300 examinees, scored by human raters and an LLM under multiple prompting conditions. For feedback, 26 rubric codes are mapped into five theory-grounded themes and modeled via dichotomous MFRM for theme detection and partial-credit MFRM for code-level severity. Results show that the surface agreement was improved with few-shot prompting, but systematic rater differences persist in where score and feedback attention are allocated and how severity is applied across themes based on a stable rater-by-domain interaction effect, indicating structured divergence. These findings demonstrate why agreement is insufficient and illustrate how a score and feedback-focused MFRM can identify systematic discrepancies that matter for assessment result interpretation.

Authors

Institutions

Publication Details

Journal
The Journal of Experimental Education
Published
2026-09-20
DOI
https://doi.org/10.1080/00220973.2026.2730125
Primary Topic
Topic Modeling
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Evaluating Theme-Specific Feedback Emphasis and Ratings in Human and Large Language Model Raters

Jiawei Xiong, George Engelhard, Cheng Tang, Jing Li
The Journal of Experimental Education
Topic Modeling
article

Evaluating Theme-Specific Feedback Emphasis and Ratings in Human and Large Language Model Raters

Jiawei Xiong, George Engelhard, Cheng Tang, Jing Li
article en

Abstract

Large Language Models (LLMs) have shown promise in automated writing evaluation that provides estimates of writing proficiency ratings and feedback, yet validation studies concentrate on score agreement, overlooking whether human and LLM raters attend to the same aspects in rubrics. We evaluate comparability between a human reference benchmark and an LLM rater using both agreement indices and Many-Facet Rasch Measurement (MFRM) in ratings and feedback, and propose the psychometric approach that separates what theme the rater emphasizes in feedback and how strongly the rater evaluates within a theme. We use an argumentative writing assessment with 300 examinees, scored by human raters and an LLM under multiple prompting conditions. For feedback, 26 rubric codes are mapped into five theory-grounded themes and modeled via dichotomous MFRM for theme detection and partial-credit MFRM for code-level severity. Results show that the surface agreement was improved with few-shot prompting, but systematic rater differences persist in where score and feedback attention are allocated and how severity is applied across themes based on a stable rater-by-domain interaction effect, indicating structured divergence. These findings demonstrate why agreement is insufficient and illustrate how a score and feedback-focused MFRM can identify systematic discrepancies that matter for assessment result interpretation.

The Journal of Experimental Education
Pittsburg State University (US)
Quality Education
Openalex Percentile: Top 8%
Topic Modeling
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Evaluating Theme-Specific Feedback Emphasis and Ratings in Human and Large Language Model Raters — Jiawei Xiong, George Engelhard, et al. · The Journal of Experimental Education (2026) | TGRS Research Map | TGRS