Development of noise resilient ASR framework for isolated Kannada word recognition
Abstract Automatic speech recognition (ASR) performance can deteriorate considerably when speech is recorded under uncontrolled and noisy acoustic conditions, particularly for low-resource languages such as Kannada. This study presents a noise-robust ASR framework for isolated Kannada speech by integrating a previously developed speech enhancement approach with acoustic modeling using the Kaldi toolkit. A Kannada speech database recorded under uncontrolled environmental conditions was transcribed and processed using a language-specific lexicon and phoneme set, followed by enhancement of the degraded speech signals prior to ASR modeling. The dataset was divided into training and testing subsets, and multiple acoustic modeling configurations were developed to investigate the effect of speech enhancement on recognition performance. The proposed enhancement framework consistently reduced the word error rate (WER) across all 16 evaluated acoustic modeling configurations. The mean WER decreased from 10.15% before enhancement to 8.73% after enhancement, corresponding to an absolute reduction of 1.42 percentage points and an average relative reduction of 14.05%. Statistical analysis using paired parametric and non-parametric tests confirmed that the observed improvement was statistically significant at the model-configuration level, with a paired $$t$$ -test yielding $$p=1.62\times10^{-6}$$ and a Wilcoxon signed-rank test yielding $$p=1.53\times10^{-5}$$ . The DNN-based configuration achieved the lowest WER, decreasing from 4.45% for noisy speech to 4.30% after enhancement. Computational analysis further indicates that the DNN acoustic model contains approximately $$8.47\times10^{5}$$ trainable parameters, requiring about 3.23 MB for model parameters in 32-bit representation. These findings demonstrate that the integration of speech enhancement with acoustic modeling provides a consistent improvement in Kannada ASR performance. The audio data collection and system implementation were carried out as part of this research. The complete Docker image, source code, and dataset are publicly available for reproducibility at https://tinyurl.com/DNNKANNADA .
Authors
- Thimmaraja Yadava G (ORCID: https://orcid.org/0000-0002-3266-9732)
- B. Dinesh Rao
- C. B. Chandrakala
- G. P. Raghudathesh
Institutions
- Nitte University (IN)
- Manipal Academy of Higher Education (IN)
Publication Details
- Journal
- International Journal of Speech Technology
- Published
- 2026-09-30
- DOI
- https://doi.org/10.1007/s10772-026-10293-6
- Primary Topic
- Speech Recognition and Synthesis
- Type
- article
- Field-Weighted Citation Impact
- 0.00