Deep learning-driven joint blind source separation and noise suppression for audio signals: a performance optimization study

Abstract Separating overlapping talkers and suppressing non-stationary noise are coupled problems, yet conventional pipelines solve them in isolation, so separation artifacts propagate untouched into the denoising stage. We therefore build a time–frequency joint separation–denoising network: a learnable encoder, a multi-scale dilated separator carrying channel and time–frequency attention, a per-stream noise-suppression module conditioned on the shared encoding, and an inverse-transform reconstructor, all reading one representation and trained end to end. A three-term objective binds scale-invariant signal-to-noise ratio, spectral magnitude and perceptual weighting, and training spans input signal-to-noise ratios down to − 5 dB. Evaluation draws on LibriMix, WHAM! and WHAMR! for separation, VoiceBank-DEMAND and MUSAN for noise, and MUSDB18-HQ for a full-band check outside speech. Against seven baselines retrained under an identical recipe, with every difference tested on 3000 mixtures across three seeds, the model reached 15.60 ± 0.10 dB SI-SNR improvement, 3.02 PESQ and 0.917 STOI at 0 dB with 3.9 million parameters. It surpassed Conv-TasNet, DPRNN, SepFormer, a diffusion-based cascade and a serial separate-then-denoise pipeline; TF-GridNet and MossFormer2 scored 0.50 to 0.80 dB higher at seven to eleven times the arithmetic cost, so the contribution is accuracy per unit of computation rather than absolute accuracy. Cost measurements report 6.8 G multiply–accumulate operations per second of audio, a real-time factor of 0.51 on a single CPU core and 0.41 J per second of processed audio, and a causal variant suited to streaming gives up 1.70 dB. Ablation separates the two attention paths, at 4.50% and 7.10%, from the joint denoising route at 10.30%. Cross-corpus, unseen-language and reverberant tests bound generalization: transfer costs 0.70 to 2.20 dB, and improvement falls to 7.80 dB once reverberation time exceeds 0.90 s. Code, configurations and per-utterance records accompany the article.

Authors

Publication Details

Journal
Scientific Reports
Published
2026-10-08
DOI
https://doi.org/10.1038/s41598-026-73211-5
Primary Topic
Speech and Audio Processing
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Deep learning-driven joint blind source separation and noise suppression for audio signals: a performance optimization study

胡延桂, Zengjie Tao, Rong Xie, Lin Lei
Scientific Reports
Speech and Audio Processing
article

Deep learning-driven joint blind source separation and noise suppression for audio signals: a performance optimization study

胡延桂, Zengjie Tao, Rong Xie, Lin Lei
article en

Abstract

Abstract Separating overlapping talkers and suppressing non-stationary noise are coupled problems, yet conventional pipelines solve them in isolation, so separation artifacts propagate untouched into the denoising stage. We therefore build a time–frequency joint separation–denoising network: a learnable encoder, a multi-scale dilated separator carrying channel and time–frequency attention, a per-stream noise-suppression module conditioned on the shared encoding, and an inverse-transform reconstructor, all reading one representation and trained end to end. A three-term objective binds scale-invariant signal-to-noise ratio, spectral magnitude and perceptual weighting, and training spans input signal-to-noise ratios down to − 5 dB. Evaluation draws on LibriMix, WHAM! and WHAMR! for separation, VoiceBank-DEMAND and MUSAN for noise, and MUSDB18-HQ for a full-band check outside speech. Against seven baselines retrained under an identical recipe, with every difference tested on 3000 mixtures across three seeds, the model reached 15.60 ± 0.10 dB SI-SNR improvement, 3.02 PESQ and 0.917 STOI at 0 dB with 3.9 million parameters. It surpassed Conv-TasNet, DPRNN, SepFormer, a diffusion-based cascade and a serial separate-then-denoise pipeline; TF-GridNet and MossFormer2 scored 0.50 to 0.80 dB higher at seven to eleven times the arithmetic cost, so the contribution is accuracy per unit of computation rather than absolute accuracy. Cost measurements report 6.8 G multiply–accumulate operations per second of audio, a real-time factor of 0.51 on a single CPU core and 0.41 J per second of processed audio, and a causal variant suited to streaming gives up 1.70 dB. Ablation separates the two attention paths, at 4.50% and 7.10%, from the joint denoising route at 10.30%. Cross-corpus, unseen-language and reverberant tests bound generalization: transfer costs 0.70 to 2.20 dB, and improvement falls to 7.80 dB once reverberation time exceeds 0.90 s. Code, configurations and per-utterance records accompany the article.

Scientific Reports
Openalex Percentile: Top 12%
Speech and Audio Processing
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.