Multi-Task Flow Matching Speech Enhancement Method with Dual Guidance of Semantic Understanding and Speaker Perception
Generative speech enhancement models often suffer from the “generative hallucination” issue, where artifacts are introduced or original speech characteristics are distorted during reconstruction due to insufficient prior constraints, severely limiting their practicality. To address this, we propose a multi-task flow-matching speech enhancement network with dual guidance from semantic understanding and speaker perception. The method employs a flow-matching model as the generative backbone, leveraging both semantic and speaker embeddings as auxiliary conditions to guide the denoising process. The semantic embedding, extracted by a pretrained audio foundation model, provides high-level content priors to ensure the intelligibility and clarity of the enhanced speech. The speaker embedding, derived from a speaker verification model, captures timbre representations, which are further regularized by a contrastive InfoNCE loss to enhance intra-class compactness and inter-class separability of speaker characteristics. During optimization, the primary denoising reconstruction task and the auxiliary speaker discrimination task are jointly optimized. The two guidance streams, together with acoustic features, are collaboratively fed into the generation network, effectively suppressing hallucinations from both content correctness and timbre consistency perspectives. Experimental results demonstrate that the proposed multi-task framework substantially mitigates hallucination-related errors in generative scenarios, achieving comprehensive improvements in objective metrics, including both reference-based measures and non-intrusive perceptual metrics, thus validating the effectiveness of the dual-guidance strategy for generative speech enhancement.
Authors
- Yongqiang Xie (ORCID: https://orcid.org/0009-0004-1857-6611)
- Xia Zou
- Xiongwei Zhang
- Chong Jia
- Yihao Li
- Meng Sun
- Bingyu Dai
Institutions
- Chinese People's Liberation Army (CN)
- PLA Army Engineering University (CN)
Publication Details
- Journal
- Electronics
- Published
- 2026-10-09
- DOI
- https://doi.org/10.3390/electronics15204607
- Primary Topic
- Speech and Audio Processing
- Type
- article
- Field-Weighted Citation Impact
- 0.00