A multimodal deep learning framework for automated oral English fluency assessment
The rapid growth in the number of oral English assessments, together with increasing demands for high accuracy and large-scale evaluation, has exposed the limitations of conventional manual scoring systems, which are often subjective, time-consuming, and labor-intensive. This study proposes a multimodal deep learning framework that integrates speech and text features for automated oral English fluency assessment. The model utilizes the Speech Content, Fluency, and Pronunciation Scores dataset obtained from Kaggle, comprising speech recordings, corresponding transcripts, and proficiency annotations, which were divided into training and testing subsets for model development and evaluation. Log-Mel spectrograms are extracted to represent acoustic characteristics, while BERT embeddings capture contextual semantic information from speech transcripts. A self-attention mechanism effectively fuses these multimodal features to evaluate fluency, pronunciation, rhythm, and emotional expression. The framework achieved 97.53% accuracy, 97.68% precision, 97.53% recall, and 97.52% F1-score. Experimental performance was evaluated against conventional single-modal deep learning approaches, demonstrating improved assessment accuracy and robustness. TensorFlow was employed for model training and optimization, whereas PyTorch was utilized for implementing and evaluating transformer-based BERT embeddings, leveraging the strengths of both frameworks within the multimodal pipeline. Integrated structure provides an efficient, scalable, and objective solution for automated oral English fluency assessment across diverse educational environments. The proposed framework incorporates audio data augmentation and ANOVA-based feature selection to improve robustness against accent variations and recording noise.
Authors
- Ping Zhang
Institutions
- Shanghai University of Political Science and Law (CN)
Publication Details
- Journal
- Discover Artificial Intelligence
- Published
- 2026-10-07
- DOI
- https://doi.org/10.1007/s44163-026-02325-6
- Primary Topic
- Speech Recognition and Synthesis
- Type
- article
- Field-Weighted Citation Impact
- 0.00