Development of noise resilient ASR framework for isolated Kannada word recognition

Abstract Automatic speech recognition (ASR) performance can deteriorate considerably when speech is recorded under uncontrolled and noisy acoustic conditions, particularly for low-resource languages such as Kannada. This study presents a noise-robust ASR framework for isolated Kannada speech by integrating a previously developed speech enhancement approach with acoustic modeling using the Kaldi toolkit. A Kannada speech database recorded under uncontrolled environmental conditions was transcribed and processed using a language-specific lexicon and phoneme set, followed by enhancement of the degraded speech signals prior to ASR modeling. The dataset was divided into training and testing subsets, and multiple acoustic modeling configurations were developed to investigate the effect of speech enhancement on recognition performance. The proposed enhancement framework consistently reduced the word error rate (WER) across all 16 evaluated acoustic modeling configurations. The mean WER decreased from 10.15% before enhancement to 8.73% after enhancement, corresponding to an absolute reduction of 1.42 percentage points and an average relative reduction of 14.05%. Statistical analysis using paired parametric and non-parametric tests confirmed that the observed improvement was statistically significant at the model-configuration level, with a paired $$t$$ -test yielding $$p=1.62\times10^{-6}$$ and a Wilcoxon signed-rank test yielding $$p=1.53\times10^{-5}$$ . The DNN-based configuration achieved the lowest WER, decreasing from 4.45% for noisy speech to 4.30% after enhancement. Computational analysis further indicates that the DNN acoustic model contains approximately $$8.47\times10^{5}$$ trainable parameters, requiring about 3.23 MB for model parameters in 32-bit representation. These findings demonstrate that the integration of speech enhancement with acoustic modeling provides a consistent improvement in Kannada ASR performance. The audio data collection and system implementation were carried out as part of this research. The complete Docker image, source code, and dataset are publicly available for reproducibility at https://tinyurl.com/DNNKANNADA .

Authors

Institutions

Publication Details

Journal
International Journal of Speech Technology
Published
2026-09-30
DOI
https://doi.org/10.1007/s10772-026-10293-6
Primary Topic
Speech Recognition and Synthesis
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Development of noise resilient ASR framework for isolated Kannada word recognition

Thimmaraja Yadava G, B. Dinesh Rao, C. B. Chandrakala, G. P. Raghudathesh
International Journal of Speech Technology
Speech Recognition and Synthesis
article

Development of noise resilient ASR framework for isolated Kannada word recognition

Thimmaraja Yadava G, B. Dinesh Rao, C. B. Chandrakala, G. P. Raghudathesh
article en

Abstract

Abstract Automatic speech recognition (ASR) performance can deteriorate considerably when speech is recorded under uncontrolled and noisy acoustic conditions, particularly for low-resource languages such as Kannada. This study presents a noise-robust ASR framework for isolated Kannada speech by integrating a previously developed speech enhancement approach with acoustic modeling using the Kaldi toolkit. A Kannada speech database recorded under uncontrolled environmental conditions was transcribed and processed using a language-specific lexicon and phoneme set, followed by enhancement of the degraded speech signals prior to ASR modeling. The dataset was divided into training and testing subsets, and multiple acoustic modeling configurations were developed to investigate the effect of speech enhancement on recognition performance. The proposed enhancement framework consistently reduced the word error rate (WER) across all 16 evaluated acoustic modeling configurations. The mean WER decreased from 10.15% before enhancement to 8.73% after enhancement, corresponding to an absolute reduction of 1.42 percentage points and an average relative reduction of 14.05%. Statistical analysis using paired parametric and non-parametric tests confirmed that the observed improvement was statistically significant at the model-configuration level, with a paired $$t$$ -test yielding $$p=1.62\times10^{-6}$$ and a Wilcoxon signed-rank test yielding $$p=1.53\times10^{-5}$$ . The DNN-based configuration achieved the lowest WER, decreasing from 4.45% for noisy speech to 4.30% after enhancement. Computational analysis further indicates that the DNN acoustic model contains approximately $$8.47\times10^{5}$$ trainable parameters, requiring about 3.23 MB for model parameters in 32-bit representation. These findings demonstrate that the integration of speech enhancement with acoustic modeling provides a consistent improvement in Kannada ASR performance. The audio data collection and system implementation were carried out as part of this research. The complete Docker image, source code, and dataset are publicly available for reproducibility at https://tinyurl.com/DNNKANNADA .

International Journal of Speech TechnologyVol. 29(4)
Nitte University (IN), Manipal Academy of Higher Education (IN)
Openalex Percentile: Top 9%
Speech Recognition and Synthesis
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.