A multimodal deep learning framework for automated oral English fluency assessment

The rapid growth in the number of oral English assessments, together with increasing demands for high accuracy and large-scale evaluation, has exposed the limitations of conventional manual scoring systems, which are often subjective, time-consuming, and labor-intensive. This study proposes a multimodal deep learning framework that integrates speech and text features for automated oral English fluency assessment. The model utilizes the Speech Content, Fluency, and Pronunciation Scores dataset obtained from Kaggle, comprising speech recordings, corresponding transcripts, and proficiency annotations, which were divided into training and testing subsets for model development and evaluation. Log-Mel spectrograms are extracted to represent acoustic characteristics, while BERT embeddings capture contextual semantic information from speech transcripts. A self-attention mechanism effectively fuses these multimodal features to evaluate fluency, pronunciation, rhythm, and emotional expression. The framework achieved 97.53% accuracy, 97.68% precision, 97.53% recall, and 97.52% F1-score. Experimental performance was evaluated against conventional single-modal deep learning approaches, demonstrating improved assessment accuracy and robustness. TensorFlow was employed for model training and optimization, whereas PyTorch was utilized for implementing and evaluating transformer-based BERT embeddings, leveraging the strengths of both frameworks within the multimodal pipeline. Integrated structure provides an efficient, scalable, and objective solution for automated oral English fluency assessment across diverse educational environments. The proposed framework incorporates audio data augmentation and ANOVA-based feature selection to improve robustness against accent variations and recording noise.

Authors

Institutions

Publication Details

Journal
Discover Artificial Intelligence
Published
2026-10-07
DOI
https://doi.org/10.1007/s44163-026-02325-6
Primary Topic
Speech Recognition and Synthesis
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

A multimodal deep learning framework for automated oral English fluency assessment

Ping Zhang
Discover Artificial Intelligence
Speech Recognition and Synthesis
article

A multimodal deep learning framework for automated oral English fluency assessment

Ping Zhang
article en

Abstract

The rapid growth in the number of oral English assessments, together with increasing demands for high accuracy and large-scale evaluation, has exposed the limitations of conventional manual scoring systems, which are often subjective, time-consuming, and labor-intensive. This study proposes a multimodal deep learning framework that integrates speech and text features for automated oral English fluency assessment. The model utilizes the Speech Content, Fluency, and Pronunciation Scores dataset obtained from Kaggle, comprising speech recordings, corresponding transcripts, and proficiency annotations, which were divided into training and testing subsets for model development and evaluation. Log-Mel spectrograms are extracted to represent acoustic characteristics, while BERT embeddings capture contextual semantic information from speech transcripts. A self-attention mechanism effectively fuses these multimodal features to evaluate fluency, pronunciation, rhythm, and emotional expression. The framework achieved 97.53% accuracy, 97.68% precision, 97.53% recall, and 97.52% F1-score. Experimental performance was evaluated against conventional single-modal deep learning approaches, demonstrating improved assessment accuracy and robustness. TensorFlow was employed for model training and optimization, whereas PyTorch was utilized for implementing and evaluating transformer-based BERT embeddings, leveraging the strengths of both frameworks within the multimodal pipeline. Integrated structure provides an efficient, scalable, and objective solution for automated oral English fluency assessment across diverse educational environments. The proposed framework incorporates audio data augmentation and ANOVA-based feature selection to improve robustness against accent variations and recording noise.

Discover Artificial IntelligenceVol. 6(1)
Shanghai University of Political Science and Law (CN)
Openalex Percentile: Top 12%
Speech Recognition and Synthesis
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.