Boosting AI-based speech performance assessment through pairwise comparative scoring of crowd data
Open-ended performance assessments can capture intellectual status that is difficult to measure using highly constrained response formats. However, their practical use is limited by the difficulty of obtaining reliable, scalable, and interpretable scores. Recent advances in large language models provide new opportunities for automated scoring, but the measurement quality of AI-generated scores remains an important empirical question. We evaluated AI models for complex speech-based assessment, and proposed a pairwise comparative scoring framework designed to improve reliability, while examining validity evidence through associations with working memory. Four speech tasks, each with three difficulty levels, were administered to 122 participants. We applied three AI scoring approaches: (a) individual scoring, where each speech response was assessed separately; (b) pairwise scoring with absolute scores, where paired responses were presented together and each participant received an absolute score; and (c) pairwise scoring with relative scores, where each score was converted into a within-pair difference. Human ratings provided a conventional benchmark. Working memory was measured to examine theoretically relevant validity evidence. Pairwise comparative scoring showed higher reliability than individual scoring, reaching levels comparable to human scores averaged across multiple raters, while providing theoretically consistent validity evidence. These results support AI-based comparative scoring for complex open-ended assessment.
Authors
- Guang Ouyang (ORCID: https://orcid.org/0000-0001-8939-7443)
- Yingzhe Li (ORCID: https://orcid.org/0000-0003-2657-8890)
Institutions
- University of Hong Kong (HK)
Publication Details
- Journal
- Scientific Reports
- Published
- 2026-10-07
- DOI
- https://doi.org/10.1038/s41598-026-74639-5
- Primary Topic
- Psychometric Methodologies and Testing
- Type
- article
- Field-Weighted Citation Impact
- 0.00