Target-speaker adaptation in text-to-speech synthesis: a comparison of efficient fine-tuning and zero-shot methods

Abstract Neural text-to-speech (TTS) systems can synthesize highly natural speech. A key capability is speaker adaptation, which enables speech synthesis that matches a target speaker’s voice characteristics, such as timbre, pitch, and prosody. Traditional neural approaches require extensive speaker-specific data and full retraining, making them resource-intensive. Recent target-speaker adaptation techniques fall into two broad categories: (i) fine-tuning-based few-shot methods, which adapt models using small amounts of target-speaker data, and (ii) zero-shot methods, which generalize to unseen speakers without additional training, using a single reference sample during inference. This work compares these two adaptation approaches, along with a training-from-scratch baseline, across different metrics. We show that fine-tuning a non-autoregressive architecture, such as ForwardTacotron, achieves speaker-adaptation quality comparable to zero-shot methods, with lower computational complexity but a lower mean opinion score. As a second contribution, we show that fine-tuning can be made more efficient in terms of data and while largely retaining quality.

Authors

Institutions

Publication Details

Journal
EURASIP Journal on Audio Speech and Music Processing
Published
2026-09-24
DOI
https://doi.org/10.1186/s13636-026-00479-w
Primary Topic
Speech Recognition and Synthesis
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Target-speaker adaptation in text-to-speech synthesis: a comparison of efficient fine-tuning and zero-shot methods

Christian Dittmar, Frank Zalkow, Emanuël A. P. Habets, Nicola Pia et al.
EURASIP Journal on Audio Speech and Music Processing
Speech Recognition and Synthesis
article

Target-speaker adaptation in text-to-speech synthesis: a comparison of efficient fine-tuning and zero-shot methods

Christian Dittmar, Frank Zalkow, Emanuël A. P. Habets, Nicola Pia, Kishor Kayyar Lakshminarayana
article en

Abstract

Abstract Neural text-to-speech (TTS) systems can synthesize highly natural speech. A key capability is speaker adaptation, which enables speech synthesis that matches a target speaker’s voice characteristics, such as timbre, pitch, and prosody. Traditional neural approaches require extensive speaker-specific data and full retraining, making them resource-intensive. Recent target-speaker adaptation techniques fall into two broad categories: (i) fine-tuning-based few-shot methods, which adapt models using small amounts of target-speaker data, and (ii) zero-shot methods, which generalize to unseen speakers without additional training, using a single reference sample during inference. This work compares these two adaptation approaches, along with a training-from-scratch baseline, across different metrics. We show that fine-tuning a non-autoregressive architecture, such as ForwardTacotron, achieves speaker-adaptation quality comparable to zero-shot methods, with lower computational complexity but a lower mean opinion score. As a second contribution, we show that fine-tuning can be made more efficient in terms of data and while largely retaining quality.

EURASIP Journal on Audio Speech and Music ProcessingVol. 2026(1)
International Audio Laboratories Erlangen (DE), Fraunhofer Institute for Integrated Circuits (DE)
Quality Education
Openalex Percentile: Top 9%
Speech Recognition and Synthesis
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Target-speaker adaptation in text-to-speech synthesis: a comparison of efficient fine-tuning and zero-shot methods — Christian Dittmar, Frank Zalkow, et al. · EURASIP Journal on Audio Speech and Music Processing (2026) | TGRS Research Map | TGRS