Multi-Task Flow Matching Speech Enhancement Method with Dual Guidance of Semantic Understanding and Speaker Perception

Generative speech enhancement models often suffer from the “generative hallucination” issue, where artifacts are introduced or original speech characteristics are distorted during reconstruction due to insufficient prior constraints, severely limiting their practicality. To address this, we propose a multi-task flow-matching speech enhancement network with dual guidance from semantic understanding and speaker perception. The method employs a flow-matching model as the generative backbone, leveraging both semantic and speaker embeddings as auxiliary conditions to guide the denoising process. The semantic embedding, extracted by a pretrained audio foundation model, provides high-level content priors to ensure the intelligibility and clarity of the enhanced speech. The speaker embedding, derived from a speaker verification model, captures timbre representations, which are further regularized by a contrastive InfoNCE loss to enhance intra-class compactness and inter-class separability of speaker characteristics. During optimization, the primary denoising reconstruction task and the auxiliary speaker discrimination task are jointly optimized. The two guidance streams, together with acoustic features, are collaboratively fed into the generation network, effectively suppressing hallucinations from both content correctness and timbre consistency perspectives. Experimental results demonstrate that the proposed multi-task framework substantially mitigates hallucination-related errors in generative scenarios, achieving comprehensive improvements in objective metrics, including both reference-based measures and non-intrusive perceptual metrics, thus validating the effectiveness of the dual-guidance strategy for generative speech enhancement.

Authors

Institutions

Publication Details

Journal
Electronics
Published
2026-10-09
DOI
https://doi.org/10.3390/electronics15204607
Primary Topic
Speech and Audio Processing
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Multi-Task Flow Matching Speech Enhancement Method with Dual Guidance of Semantic Understanding and Speaker Perception

Yongqiang Xie, Xia Zou, Xiongwei Zhang, Chong Jia et al.
Electronics
Speech and Audio Processing
article

Multi-Task Flow Matching Speech Enhancement Method with Dual Guidance of Semantic Understanding and Speaker Perception

Yongqiang Xie, Xia Zou, Xiongwei Zhang, Chong Jia, Yihao Li, Meng Sun, Bingyu Dai
article en

Abstract

Generative speech enhancement models often suffer from the “generative hallucination” issue, where artifacts are introduced or original speech characteristics are distorted during reconstruction due to insufficient prior constraints, severely limiting their practicality. To address this, we propose a multi-task flow-matching speech enhancement network with dual guidance from semantic understanding and speaker perception. The method employs a flow-matching model as the generative backbone, leveraging both semantic and speaker embeddings as auxiliary conditions to guide the denoising process. The semantic embedding, extracted by a pretrained audio foundation model, provides high-level content priors to ensure the intelligibility and clarity of the enhanced speech. The speaker embedding, derived from a speaker verification model, captures timbre representations, which are further regularized by a contrastive InfoNCE loss to enhance intra-class compactness and inter-class separability of speaker characteristics. During optimization, the primary denoising reconstruction task and the auxiliary speaker discrimination task are jointly optimized. The two guidance streams, together with acoustic features, are collaboratively fed into the generation network, effectively suppressing hallucinations from both content correctness and timbre consistency perspectives. Experimental results demonstrate that the proposed multi-task framework substantially mitigates hallucination-related errors in generative scenarios, achieving comprehensive improvements in objective metrics, including both reference-based measures and non-intrusive perceptual metrics, thus validating the effectiveness of the dual-guidance strategy for generative speech enhancement.

ElectronicsVol. 15(20)
Chinese People's Liberation Army (CN), PLA Army Engineering University (CN)
Openalex Percentile: Top 12%
Speech and Audio Processing
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Multi-Task Flow Matching Speech Enhancement Method with Dual Guidance of Semantic Understanding and Speaker Perception — Yongqiang Xie, Xia Zou, et al. · Electronics (2026) | TGRS Research Map | TGRS