Automatic disfluency correction for a linguistically complex language

Abstract Automatic disfluency correction is an important task in natural language processing, particularly for morphologically rich and low-resource languages such as Tamil. Disfluencies including repetitions, interjections, prolongations, and part-word repetitions can negatively affect the performance of speech transcription systems, conversational agents, and downstream language processing applications. This paper presents a hybrid BERT-BiLSTM-Attention framework for token-level Tamil disfluency correction using binary KEEP and DELETE sequence labeling. The proposed model integrates contextual multilingual BERT embeddings with bidirectional sequential modeling and an attention mechanism for token-level disfluency correction. To address the limited availability of Tamil disfluency correction resources, a Synthetic Tamil Disfluent Text Corpus (STDTC) was developed using rule-based augmentation strategies. Experiments were conducted on both the synthetic STDTC dataset and a merged evaluation setting incorporating a small curated subset of SPRING_INX_Tamil_R2 data. The proposed framework achieved a token-level accuracy of 96.57% and a macro F1-score of 96.46% on the STDTC dataset, and 96.77% accuracy and 96.67% macro F1-score on the merged evaluation setting (STDTC + SPRING_INX_Tamil_R2). This merged setting reflects a primarily synthetic evaluation supplemented by a small number of real conversational examples. Additional analyses, including deletion-specific precision and recall, confusion matrices, exact sentence match analysis, and bootstrap confidence intervals, were performed to assess model performance across multiple disfluency categories. The results indicate that the proposed framework can model common Tamil disfluency patterns while maintaining consistent token-level correction performance. The study further suggests that combining contextual embeddings, sequential learning, and attention mechanisms is beneficial for low-resource Tamil disfluency correction tasks.

Authors

Publication Details

Journal
Scientific Reports
Published
2026-10-08
DOI
https://doi.org/10.1038/s41598-026-68238-7
Primary Topic
Speech Recognition and Synthesis
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Automatic disfluency correction for a linguistically complex language

Sheena Christabel Pravin, Rajasekar M.
Scientific Reports
Speech Recognition and Synthesis
article

Automatic disfluency correction for a linguistically complex language

Sheena Christabel Pravin, Rajasekar M.
article en

Abstract

Abstract Automatic disfluency correction is an important task in natural language processing, particularly for morphologically rich and low-resource languages such as Tamil. Disfluencies including repetitions, interjections, prolongations, and part-word repetitions can negatively affect the performance of speech transcription systems, conversational agents, and downstream language processing applications. This paper presents a hybrid BERT-BiLSTM-Attention framework for token-level Tamil disfluency correction using binary KEEP and DELETE sequence labeling. The proposed model integrates contextual multilingual BERT embeddings with bidirectional sequential modeling and an attention mechanism for token-level disfluency correction. To address the limited availability of Tamil disfluency correction resources, a Synthetic Tamil Disfluent Text Corpus (STDTC) was developed using rule-based augmentation strategies. Experiments were conducted on both the synthetic STDTC dataset and a merged evaluation setting incorporating a small curated subset of SPRING_INX_Tamil_R2 data. The proposed framework achieved a token-level accuracy of 96.57% and a macro F1-score of 96.46% on the STDTC dataset, and 96.77% accuracy and 96.67% macro F1-score on the merged evaluation setting (STDTC + SPRING_INX_Tamil_R2). This merged setting reflects a primarily synthetic evaluation supplemented by a small number of real conversational examples. Additional analyses, including deletion-specific precision and recall, confusion matrices, exact sentence match analysis, and bootstrap confidence intervals, were performed to assess model performance across multiple disfluency categories. The results indicate that the proposed framework can model common Tamil disfluency patterns while maintaining consistent token-level correction performance. The study further suggests that combining contextual embeddings, sequential learning, and attention mechanisms is beneficial for low-resource Tamil disfluency correction tasks.

Scientific Reports
Openalex Percentile: Top 12%
Speech Recognition and Synthesis
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Automatic disfluency correction for a linguistically complex language — Sheena Christabel Pravin, Rajasekar M. · Scientific Reports (2026) | TGRS Research Map | TGRS