Evaluating Theme-Specific Feedback Emphasis and Ratings in Human and Large Language Model Raters
Large Language Models (LLMs) have shown promise in automated writing evaluation that provides estimates of writing proficiency ratings and feedback, yet validation studies concentrate on score agreement, overlooking whether human and LLM raters attend to the same aspects in rubrics. We evaluate comparability between a human reference benchmark and an LLM rater using both agreement indices and Many-Facet Rasch Measurement (MFRM) in ratings and feedback, and propose the psychometric approach that separates what theme the rater emphasizes in feedback and how strongly the rater evaluates within a theme. We use an argumentative writing assessment with 300 examinees, scored by human raters and an LLM under multiple prompting conditions. For feedback, 26 rubric codes are mapped into five theory-grounded themes and modeled via dichotomous MFRM for theme detection and partial-credit MFRM for code-level severity. Results show that the surface agreement was improved with few-shot prompting, but systematic rater differences persist in where score and feedback attention are allocated and how severity is applied across themes based on a stable rater-by-domain interaction effect, indicating structured divergence. These findings demonstrate why agreement is insufficient and illustrate how a score and feedback-focused MFRM can identify systematic discrepancies that matter for assessment result interpretation.
Authors
- Jiawei Xiong (ORCID: https://orcid.org/0000-0002-2069-8720)
- George Engelhard (ORCID: https://orcid.org/0000-0002-1694-8942)
- Cheng Tang
- Jing Li
Institutions
- Pittsburg State University (US)
Publication Details
- Journal
- The Journal of Experimental Education
- Published
- 2026-09-20
- DOI
- https://doi.org/10.1080/00220973.2026.2730125
- Primary Topic
- Topic Modeling
- Type
- article
- Field-Weighted Citation Impact
- 0.00