Low-Resource Vietnamese Logistics Outbound-Call Speech Recognition with Low-Rank Adaptation and Dynamic Curriculum Learning

Reliable manual annotations for Vietnamese logistics calls are scarce, and telephone channel degradation compounds errors in business-critical entities. Existing adaptation methods typically optimize aggregate transcription accuracy, use uniform token losses, or apply a fixed training schedule; few control parameter updates, entity-level supervision, and sample order together when labels are noisy. We address this problem by combining low-rank adaptation (LoRA), token-level entity-weighted cross-entropy, and validation feedback curriculum learning for Whisper-large-v3. A joint acoustic–semantic score ranks samples by signal quality, sequence length, entity density, and linguistic rarity. The curriculum boundary expands only after the validation loss improves. On 20,317 in-domain training utterances, the complete system achieves a word error rate (WER) of 0.0735 and a mean exact match entity accuracy of 92.1%. Compared with vanilla LoRA, it reduces WER by 4.92% and raises entity accuracy from 85.3% to 92.1%. Ablation results show that curriculum pacing and entity weighting each contribute to the final performance.

Authors

Institutions

Publication Details

Journal
Big Data and Cognitive Computing
Published
2026-10-07
DOI
https://doi.org/10.3390/bdcc10100341
Primary Topic
Speech Recognition and Synthesis
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Low-Resource Vietnamese Logistics Outbound-Call Speech Recognition with Low-Rank Adaptation and Dynamic Curriculum Learning

Wenjiang Ji, Zhiqiang Zhao, Yuan Qiu, Xiaoxiong Zhang
Big Data and Cognitive Computing
Speech Recognition and Synthesis
article

Low-Resource Vietnamese Logistics Outbound-Call Speech Recognition with Low-Rank Adaptation and Dynamic Curriculum Learning

Wenjiang Ji, Zhiqiang Zhao, Yuan Qiu, Xiaoxiong Zhang
article en

Abstract

Reliable manual annotations for Vietnamese logistics calls are scarce, and telephone channel degradation compounds errors in business-critical entities. Existing adaptation methods typically optimize aggregate transcription accuracy, use uniform token losses, or apply a fixed training schedule; few control parameter updates, entity-level supervision, and sample order together when labels are noisy. We address this problem by combining low-rank adaptation (LoRA), token-level entity-weighted cross-entropy, and validation feedback curriculum learning for Whisper-large-v3. A joint acoustic–semantic score ranks samples by signal quality, sequence length, entity density, and linguistic rarity. The curriculum boundary expands only after the validation loss improves. On 20,317 in-domain training utterances, the complete system achieves a word error rate (WER) of 0.0735 and a mean exact match entity accuracy of 92.1%. Compared with vanilla LoRA, it reduces WER by 4.92% and raises entity accuracy from 85.3% to 92.1%. Ablation results show that curriculum pacing and entity weighting each contribute to the final performance.

Big Data and Cognitive ComputingVol. 10(10)
Xi'an University of Technology (CN)
Openalex Percentile: Top 12%
Speech Recognition and Synthesis
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Low-Resource Vietnamese Logistics Outbound-Call Speech Recognition with Low-Rank Adaptation and Dynamic Curriculum Learning — Wenjiang Ji, Zhiqiang Zhao, et al. · Big Data and Cognitive Computing (2026) | TGRS Research Map | TGRS