Automatic disfluency correction for a linguistically complex language
Abstract Automatic disfluency correction is an important task in natural language processing, particularly for morphologically rich and low-resource languages such as Tamil. Disfluencies including repetitions, interjections, prolongations, and part-word repetitions can negatively affect the performance of speech transcription systems, conversational agents, and downstream language processing applications. This paper presents a hybrid BERT-BiLSTM-Attention framework for token-level Tamil disfluency correction using binary KEEP and DELETE sequence labeling. The proposed model integrates contextual multilingual BERT embeddings with bidirectional sequential modeling and an attention mechanism for token-level disfluency correction. To address the limited availability of Tamil disfluency correction resources, a Synthetic Tamil Disfluent Text Corpus (STDTC) was developed using rule-based augmentation strategies. Experiments were conducted on both the synthetic STDTC dataset and a merged evaluation setting incorporating a small curated subset of SPRING_INX_Tamil_R2 data. The proposed framework achieved a token-level accuracy of 96.57% and a macro F1-score of 96.46% on the STDTC dataset, and 96.77% accuracy and 96.67% macro F1-score on the merged evaluation setting (STDTC + SPRING_INX_Tamil_R2). This merged setting reflects a primarily synthetic evaluation supplemented by a small number of real conversational examples. Additional analyses, including deletion-specific precision and recall, confusion matrices, exact sentence match analysis, and bootstrap confidence intervals, were performed to assess model performance across multiple disfluency categories. The results indicate that the proposed framework can model common Tamil disfluency patterns while maintaining consistent token-level correction performance. The study further suggests that combining contextual embeddings, sequential learning, and attention mechanisms is beneficial for low-resource Tamil disfluency correction tasks.
Authors
- Sheena Christabel Pravin (ORCID: https://orcid.org/0000-0001-8520-3322)
- Rajasekar M.
Publication Details
- Journal
- Scientific Reports
- Published
- 2026-10-08
- DOI
- https://doi.org/10.1038/s41598-026-68238-7
- Primary Topic
- Speech Recognition and Synthesis
- Type
- article
- Field-Weighted Citation Impact
- 0.00