Low-Resource Vietnamese Logistics Outbound-Call Speech Recognition with Low-Rank Adaptation and Dynamic Curriculum Learning
Reliable manual annotations for Vietnamese logistics calls are scarce, and telephone channel degradation compounds errors in business-critical entities. Existing adaptation methods typically optimize aggregate transcription accuracy, use uniform token losses, or apply a fixed training schedule; few control parameter updates, entity-level supervision, and sample order together when labels are noisy. We address this problem by combining low-rank adaptation (LoRA), token-level entity-weighted cross-entropy, and validation feedback curriculum learning for Whisper-large-v3. A joint acoustic–semantic score ranks samples by signal quality, sequence length, entity density, and linguistic rarity. The curriculum boundary expands only after the validation loss improves. On 20,317 in-domain training utterances, the complete system achieves a word error rate (WER) of 0.0735 and a mean exact match entity accuracy of 92.1%. Compared with vanilla LoRA, it reduces WER by 4.92% and raises entity accuracy from 85.3% to 92.1%. Ablation results show that curriculum pacing and entity weighting each contribute to the final performance.
Authors
- Wenjiang Ji (ORCID: https://orcid.org/0000-0001-5502-1579)
- Zhiqiang Zhao (ORCID: https://orcid.org/0000-0002-2475-7177)
- Yuan Qiu (ORCID: https://orcid.org/0009-0005-0850-4777)
- Xiaoxiong Zhang (ORCID: https://orcid.org/0000-0002-8801-8162)
Institutions
- Xi'an University of Technology (CN)
Publication Details
- Journal
- Big Data and Cognitive Computing
- Published
- 2026-10-07
- DOI
- https://doi.org/10.3390/bdcc10100341
- Primary Topic
- Speech Recognition and Synthesis
- Type
- article
- Field-Weighted Citation Impact
- 0.00