Enhancing Quranic Learning: A Multimodal Deep Learning Approach for Arabic Phoneme Recognition
Arabic phoneme recognition is relevant to computer-assisted Qur’anic learning, but short isolated recordings provide limited context and subtle acoustic distinctions remain difficult to model. This study presents an intermediate-fusion framework combining UniSpeech acoustic representations with BERT embeddings of Whisper-generated text. Early and staged late fusion are evaluated as alternative integration strategies using two configurations of an existing 1015-recording corpus with 29 target classes. Dataset A uses 783 recordings for training and 232 for testing, whereas Dataset B uses 812 and 203 recordings, respectively. In the reported experiments, the proposed intermediate-fusion model achieves accuracies of 95.69% and 98.03% on Dataset A and Dataset B, respectively. Early fusion reaches 96.55% on Dataset A, showing that the proposed method is not uniformly the numerical leader. On Dataset B, this corresponds to 199 correct predictions, compared with 196 for early fusion and 193 for staged late fusion. Exact paired McNemar tests show no significant difference among the three strategies after Holm correction at the 0.05 level. The contribution is an empirical comparison of fusion strategies using pretrained models, rather than a new transformer architecture. The results support further investigation of audio-derived textual representations for Arabic phoneme recognition, while the small corpus, incomplete speaker metadata, and absence of a separate pronunciation-error evaluation limit conclusions about speaker-independent pronunciation assessment and practical learning outcomes.
Authors
- Ayhan Küçükmanіşa (ORCID: https://orcid.org/0000-0002-1886-1250)
- Zeynep Hilal Kilimci (ORCID: https://orcid.org/0000-0003-1497-305X)
- Şükrü Selim Çalık (ORCID: https://orcid.org/0000-0002-3884-7966)
- Derya Gelmez
Institutions
- Kocaeli Üniversitesi (TR)
Publication Details
- Journal
- Applied Sciences
- Published
- 2026-10-07
- DOI
- https://doi.org/10.3390/app16199905
- Primary Topic
- Speech Recognition and Synthesis
- Type
- article
- Field-Weighted Citation Impact
- 0.00