Boosting AI-based speech performance assessment through pairwise comparative scoring of crowd data

Open-ended performance assessments can capture intellectual status that is difficult to measure using highly constrained response formats. However, their practical use is limited by the difficulty of obtaining reliable, scalable, and interpretable scores. Recent advances in large language models provide new opportunities for automated scoring, but the measurement quality of AI-generated scores remains an important empirical question. We evaluated AI models for complex speech-based assessment, and proposed a pairwise comparative scoring framework designed to improve reliability, while examining validity evidence through associations with working memory. Four speech tasks, each with three difficulty levels, were administered to 122 participants. We applied three AI scoring approaches: (a) individual scoring, where each speech response was assessed separately; (b) pairwise scoring with absolute scores, where paired responses were presented together and each participant received an absolute score; and (c) pairwise scoring with relative scores, where each score was converted into a within-pair difference. Human ratings provided a conventional benchmark. Working memory was measured to examine theoretically relevant validity evidence. Pairwise comparative scoring showed higher reliability than individual scoring, reaching levels comparable to human scores averaged across multiple raters, while providing theoretically consistent validity evidence. These results support AI-based comparative scoring for complex open-ended assessment.

Authors

Institutions

Publication Details

Journal
Scientific Reports
Published
2026-10-07
DOI
https://doi.org/10.1038/s41598-026-74639-5
Primary Topic
Psychometric Methodologies and Testing
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Boosting AI-based speech performance assessment through pairwise comparative scoring of crowd data

Guang Ouyang, Yingzhe Li
Scientific Reports
Psychometric Methodologies and Testing
article

Boosting AI-based speech performance assessment through pairwise comparative scoring of crowd data

Guang Ouyang, Yingzhe Li
article en

Abstract

Open-ended performance assessments can capture intellectual status that is difficult to measure using highly constrained response formats. However, their practical use is limited by the difficulty of obtaining reliable, scalable, and interpretable scores. Recent advances in large language models provide new opportunities for automated scoring, but the measurement quality of AI-generated scores remains an important empirical question. We evaluated AI models for complex speech-based assessment, and proposed a pairwise comparative scoring framework designed to improve reliability, while examining validity evidence through associations with working memory. Four speech tasks, each with three difficulty levels, were administered to 122 participants. We applied three AI scoring approaches: (a) individual scoring, where each speech response was assessed separately; (b) pairwise scoring with absolute scores, where paired responses were presented together and each participant received an absolute score; and (c) pairwise scoring with relative scores, where each score was converted into a within-pair difference. Human ratings provided a conventional benchmark. Working memory was measured to examine theoretically relevant validity evidence. Pairwise comparative scoring showed higher reliability than individual scoring, reaching levels comparable to human scores averaged across multiple raters, while providing theoretically consistent validity evidence. These results support AI-based comparative scoring for complex open-ended assessment.

Scientific Reports
University of Hong Kong (HK)
Openalex Percentile: Top 9%
Psychometric Methodologies and Testing
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Boosting AI-based speech performance assessment through pairwise comparative scoring of crowd data — Guang Ouyang, Yingzhe Li · Scientific Reports (2026) | TGRS Research Map | TGRS