Enhancing Quranic Learning: A Multimodal Deep Learning Approach for Arabic Phoneme Recognition

Arabic phoneme recognition is relevant to computer-assisted Qur’anic learning, but short isolated recordings provide limited context and subtle acoustic distinctions remain difficult to model. This study presents an intermediate-fusion framework combining UniSpeech acoustic representations with BERT embeddings of Whisper-generated text. Early and staged late fusion are evaluated as alternative integration strategies using two configurations of an existing 1015-recording corpus with 29 target classes. Dataset A uses 783 recordings for training and 232 for testing, whereas Dataset B uses 812 and 203 recordings, respectively. In the reported experiments, the proposed intermediate-fusion model achieves accuracies of 95.69% and 98.03% on Dataset A and Dataset B, respectively. Early fusion reaches 96.55% on Dataset A, showing that the proposed method is not uniformly the numerical leader. On Dataset B, this corresponds to 199 correct predictions, compared with 196 for early fusion and 193 for staged late fusion. Exact paired McNemar tests show no significant difference among the three strategies after Holm correction at the 0.05 level. The contribution is an empirical comparison of fusion strategies using pretrained models, rather than a new transformer architecture. The results support further investigation of audio-derived textual representations for Arabic phoneme recognition, while the small corpus, incomplete speaker metadata, and absence of a separate pronunciation-error evaluation limit conclusions about speaker-independent pronunciation assessment and practical learning outcomes.

Authors

Institutions

Publication Details

Journal
Applied Sciences
Published
2026-10-07
DOI
https://doi.org/10.3390/app16199905
Primary Topic
Speech Recognition and Synthesis
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Enhancing Quranic Learning: A Multimodal Deep Learning Approach for Arabic Phoneme Recognition

Ayhan Küçükmanіşa, Zeynep Hilal Kilimci, Şükrü Selim Çalık, Derya Gelmez
Applied Sciences
Speech Recognition and Synthesis
article

Enhancing Quranic Learning: A Multimodal Deep Learning Approach for Arabic Phoneme Recognition

Ayhan Küçükmanіşa, Zeynep Hilal Kilimci, Şükrü Selim Çalık, Derya Gelmez
article en

Abstract

Arabic phoneme recognition is relevant to computer-assisted Qur’anic learning, but short isolated recordings provide limited context and subtle acoustic distinctions remain difficult to model. This study presents an intermediate-fusion framework combining UniSpeech acoustic representations with BERT embeddings of Whisper-generated text. Early and staged late fusion are evaluated as alternative integration strategies using two configurations of an existing 1015-recording corpus with 29 target classes. Dataset A uses 783 recordings for training and 232 for testing, whereas Dataset B uses 812 and 203 recordings, respectively. In the reported experiments, the proposed intermediate-fusion model achieves accuracies of 95.69% and 98.03% on Dataset A and Dataset B, respectively. Early fusion reaches 96.55% on Dataset A, showing that the proposed method is not uniformly the numerical leader. On Dataset B, this corresponds to 199 correct predictions, compared with 196 for early fusion and 193 for staged late fusion. Exact paired McNemar tests show no significant difference among the three strategies after Holm correction at the 0.05 level. The contribution is an empirical comparison of fusion strategies using pretrained models, rather than a new transformer architecture. The results support further investigation of audio-derived textual representations for Arabic phoneme recognition, while the small corpus, incomplete speaker metadata, and absence of a separate pronunciation-error evaluation limit conclusions about speaker-independent pronunciation assessment and practical learning outcomes.

Applied SciencesVol. 16(19)
Kocaeli Üniversitesi (TR)
Openalex Percentile: Top 12%
Speech Recognition and Synthesis
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Enhancing Quranic Learning: A Multimodal Deep Learning Approach for Arabic Phoneme Recognition — Ayhan Küçükmanіşa, Zeynep Hilal Kilimci, et al. · Applied Sciences (2026) | TGRS Research Map | TGRS