Speech intelligibility comparison of standalone and two-stage deep learning architectures for behind-the-ear-to-binaural enhancement

Abstract The relative perceptual performance of alternative deep learning based speech processing architectures remains uncertain under complex acoustic conditions. This study compared speech intelligibility, in listeners with normal hearing and with mild-to-moderate hearing loss, obtained with two processing models: a speech-target-oriented two-stage model (SED + U-Net), designed to estimate speech-target activity before denoising, and a standalone U-Net model. These two architectures were selected as representative instances of two contrasting design philosophies in deep-learning-based speech enhancement: unconstrained direct denoising versus denoising explicitly guided by a prior estimate of speech-target activity. Intelligibility was assessed using a keyword-recognition task with sentences from the Chilean SHARVARD corpus, presented under controlled virtual acoustic conditions. Twenty-four adults participated in the study, consisting of 12 normal-hearing (NH) listeners and 12 listeners with mild-to-moderate hearing loss (HL). Experimental conditions varied by processing model, masker configuration (Masker Around, Masker Front, and Masker Side), reverberation time (RT = 0.2 s and 0.8 s), and signal-to-noise ratio (− 3, 0, 3, 6, and 9 dB). Keyword recognition was analyzed using a binomial generalized linear mixed-effects model. Results showed significant effects of SNR, processing model, masker configuration, and reverberation time. Unlike the original hypothesis, the standalone U-Net performed better in keyword recognition than the speech-target-oriented two-stage model, emphasizing the need for perceptual validation in evaluating speech-processing systems.

Authors

Institutions

Publication Details

Journal
Scientific Reports
Published
2026-09-18
DOI
https://doi.org/10.1038/s41598-026-72055-3
Primary Topic
Hearing Loss and Rehabilitation
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Speech intelligibility comparison of standalone and two-stage deep learning architectures for behind-the-ear-to-binaural enhancement

Rhoddy Viveros-Muñoz, Jorge P. Arenas, Carla E. Contreras-Saavedra, Sebastián Guajardo Herrera
Scientific Reports
Hearing Loss and Rehabilitation
article

Speech intelligibility comparison of standalone and two-stage deep learning architectures for behind-the-ear-to-binaural enhancement

Rhoddy Viveros-Muñoz, Jorge P. Arenas, Carla E. Contreras-Saavedra, Sebastián Guajardo Herrera
article en

Abstract

Abstract The relative perceptual performance of alternative deep learning based speech processing architectures remains uncertain under complex acoustic conditions. This study compared speech intelligibility, in listeners with normal hearing and with mild-to-moderate hearing loss, obtained with two processing models: a speech-target-oriented two-stage model (SED + U-Net), designed to estimate speech-target activity before denoising, and a standalone U-Net model. These two architectures were selected as representative instances of two contrasting design philosophies in deep-learning-based speech enhancement: unconstrained direct denoising versus denoising explicitly guided by a prior estimate of speech-target activity. Intelligibility was assessed using a keyword-recognition task with sentences from the Chilean SHARVARD corpus, presented under controlled virtual acoustic conditions. Twenty-four adults participated in the study, consisting of 12 normal-hearing (NH) listeners and 12 listeners with mild-to-moderate hearing loss (HL). Experimental conditions varied by processing model, masker configuration (Masker Around, Masker Front, and Masker Side), reverberation time (RT = 0.2 s and 0.8 s), and signal-to-noise ratio (− 3, 0, 3, 6, and 9 dB). Keyword recognition was analyzed using a binomial generalized linear mixed-effects model. Results showed significant effects of SNR, processing model, masker configuration, and reverberation time. Unlike the original hypothesis, the standalone U-Net performed better in keyword recognition than the speech-target-oriented two-stage model, emphasizing the need for perceptual validation in evaluating speech-processing systems.

Scientific Reports
Austral University of Chile (CL), University of Concepción (CL), San Sebastián University (CL), Universidad Católica de la Santísima Concepción (CL), University of Bío-Bío (CL), Federico Santa María Technical University (CL)
Agencia Nacional de Investigación y Desarrollo
Openalex Percentile: Top 9%
Hearing Loss and Rehabilitation
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.