Lithuanian speech recognition: effects of dialect training and transcript spelling
Enterprise conversation intelligence depends on accurate speech recognition, including when speakers use regional dialects. We test whether adding dialect speech improves Lithuanian speech recognition and whether dialect transcripts should use dialect or standard spelling. We fine-tune Parakeet-TDT on 420.17 hours of Lithuanian telephone speech, adding up to 82.30 hours of dialect recordings. We compare this with adding other spontaneous speech and test dialect and standard transcript spelling. Training settings are fixed. We evaluate dialect recognition against both spellings and an "either-form" score that accepts either reference form. A separate 2.78-hour Customer service speech test contains previously unseen speakers. Adding dialect speech reduces word error rate (WER) against dialect references from 41.62% to 31.30%, averaged over two training runs. Adding other spontaneous speech lowers it only to 40.04%. In the single-run comparison, the first 25.43 hours provide about two thirds of the dialect-test improvement. Standard-spelling dialect transcripts give the lowest either-form WER among models also trained on telephone speech: 25.97%. On Customer service speech, adding dialect recordings lowers WER in each of two training runs, from an average of 44.89% to 39.43% with dialect spelling and to 39.06% with standard spelling; run-to-run variation on this test is larger than on the dialect test. WER on the LIEPA-3 telephone test changes little; that test may share speakers with training. These results support adding dialect recordings for regional speech and customer-service conversations. Transcript spelling affects the output learned and the accuracy measured.
Authors
- Saulius Jarasiunas
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-05
- DOI
- https://doi.org/10.5281/zenodo.23156467
- Primary Topic
- Speech Recognition and Synthesis
- Type
- preprint