Novel Approaches to Language Assessment Part 1: Artificial Intelligence–Automated Scoring of a Sentence Recall Task for School-Age Children

PURPOSE: Artificial intelligence (AI) technologies have the potential to enhance the delivery of speech-language pathology services. One area of practice that could benefit from AI is automated screenings for developmental language disorder (DLD) in early elementary students. This report presents the elements and processes involved in evaluating an AI-based automated scoring system using archival recordings. METHOD: A total of 947 audio samples of children completing the Redmond Sentence Recall (RSR) task, representing cases with and without DLD, were used. We developed an automated version of the RSR (AutoRSR), employing automated speech recognition (ASR), natural language processing, and machine learning models, to automate the tasks of transcribing recordings, tracking the number of errors children produced, scoring production accuracies based on RSR criteria, and generating pass/fail decisions based on age-referenced expectations against a reference standard (i.e., Core Language scores from the Clinical Evaluation of Language Fundamentals-Fourth Edition for a subset of 130 children). Different speech recognition models (Whisper and Reverb ASR) were applied to examine potential trade-offs. The accuracy of the AutoRSR was compared to human scoring. RESULTS: AutoRSR with Whisper speech recognition yielded sensitivity = .768 and specificity = .784, whereas Reverb speech recognition yielded sensitivity = .873 and specificity = .730. While AutoRSR had some mismatches with human evaluators and tended to produce lower scores, it converged more closely with the reference standard than human-scored protocols. CONCLUSIONS: AI technologies have become powerful enough to automate the scoring of recordings of children's sentence recall with similar accuracy levels as human evaluators. Next steps are discussed. SUPPLEMENTAL MATERIAL: https://doi.org/10.23641/asha.34050876.

Authors

Institutions

Publication Details

Journal
Language Speech and Hearing Services in Schools
Published
2026-10-09
DOI
https://doi.org/10.1044/2026_lshss-25-00268
Primary Topic
Language Development and Disorders
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Novel Approaches to Language Assessment Part 1: Artificial Intelligence–Automated Scoring of a Sentence Recall Task for School-Age Children

Dancheng Liu, Jinjun Xiong, Pamela A. Hadley, Carol Miller et al.
Language Speech and Hearing Services in Schools
Language Development and Disorders
article

Novel Approaches to Language Assessment Part 1: Artificial Intelligence–Automated Scoring of a Sentence Recall Task for School-Age Children

Dancheng Liu, Jinjun Xiong, Pamela A. Hadley, Carol Miller, Sean M. Redmond, Jason Yang
article en

Abstract

PURPOSE: Artificial intelligence (AI) technologies have the potential to enhance the delivery of speech-language pathology services. One area of practice that could benefit from AI is automated screenings for developmental language disorder (DLD) in early elementary students. This report presents the elements and processes involved in evaluating an AI-based automated scoring system using archival recordings. METHOD: A total of 947 audio samples of children completing the Redmond Sentence Recall (RSR) task, representing cases with and without DLD, were used. We developed an automated version of the RSR (AutoRSR), employing automated speech recognition (ASR), natural language processing, and machine learning models, to automate the tasks of transcribing recordings, tracking the number of errors children produced, scoring production accuracies based on RSR criteria, and generating pass/fail decisions based on age-referenced expectations against a reference standard (i.e., Core Language scores from the Clinical Evaluation of Language Fundamentals-Fourth Edition for a subset of 130 children). Different speech recognition models (Whisper and Reverb ASR) were applied to examine potential trade-offs. The accuracy of the AutoRSR was compared to human scoring. RESULTS: AutoRSR with Whisper speech recognition yielded sensitivity = .768 and specificity = .784, whereas Reverb speech recognition yielded sensitivity = .873 and specificity = .730. While AutoRSR had some mismatches with human evaluators and tended to produce lower scores, it converged more closely with the reference standard than human-scored protocols. CONCLUSIONS: AI technologies have become powerful enough to automate the scoring of recordings of children's sentence recall with similar accuracy levels as human evaluators. Next steps are discussed. SUPPLEMENTAL MATERIAL: https://doi.org/10.23641/asha.34050876.

Language Speech and Hearing Services in Schools
Pennsylvania State University (US), University of Illinois Urbana-Champaign (US), University of Utah (US), The University of Texas at San Antonio (US), University at Buffalo, State University of New York (US)
Openalex Percentile: Top 5%
Language Development and Disorders
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.