Uneven pathways from ASR evidence to oral action: calibrated uptake in a multimodal GenAI-supported English academic speaking module
Purpose: Automated speech recognition (ASR) and generative artificial intelligence (GenAI) increasingly support second language speaking, but less is known about how learners transform audio, transcript, textual feedback, task evidence, and human support into subsequent oral action. This study examined within-module speaking change and uneven uptake in a multimodal ASR-informed GenAI-supported English academic speaking module.Design/methodology/approach: This single-arm explanatory mixed-methods pre–post study involved Chinese postgraduate exchange and visiting students in a New Zealand speaking salon. Of 115 retained module records, 104 learners with valid pre- and post-module speaking scores, survey responses, platform records, and sufficient core-session participation were included in the primary quantitative analyses. The module organised practice as a teachable cycle of spoken response, ASR/replay noticing, targeted GenAI consultation, cross-source verification, selective uptake, and oral rehearsal.Findings and Originality/value: Rated speaking increased from M = 3.04 to M = 3.27, t(103) = 5.41, p < 0.001, dz = 0.53, although 66 learners improved, 9 were unchanged, and 29 declined. Reported module-contextualised proactive engagement was positively associated with post-module performance after controls (B = 0.170, p < 0.001), whereas speaking anxiety and the exploratory interaction were not significant. Platform activity showed a small positive relation with gain (r = 0.222, p = 0.024), while qualitative evidence suggested that activity became educationally consequential when learners identified specific targets, verified suggestions, adapted them, and rehearsed revised oral responses. The study conceptualises multimodality as learner-managed transformation across representations and identifies targeted noticing, verified AI feedback, and oral rehearsal as the pedagogical core of uneven uptake.
Authors
- Lawrence Jun Zhang (ORCID: https://orcid.org/0000-0003-1025-1746)
- Huican Huo (ORCID: https://orcid.org/0009-0001-7650-2950)
Institutions
- University of Auckland (NZ)
Publication Details
- Journal
- Innovation in Language Learning and Teaching
- Published
- 2026-09-16
- DOI
- https://doi.org/10.1080/17501229.2026.2729888
- Primary Topic
- AI in Service Interactions
- Type
- article
- Field-Weighted Citation Impact
- 0.00