Target-speaker adaptation in text-to-speech synthesis: a comparison of efficient fine-tuning and zero-shot methods
Abstract Neural text-to-speech (TTS) systems can synthesize highly natural speech. A key capability is speaker adaptation, which enables speech synthesis that matches a target speaker’s voice characteristics, such as timbre, pitch, and prosody. Traditional neural approaches require extensive speaker-specific data and full retraining, making them resource-intensive. Recent target-speaker adaptation techniques fall into two broad categories: (i) fine-tuning-based few-shot methods, which adapt models using small amounts of target-speaker data, and (ii) zero-shot methods, which generalize to unseen speakers without additional training, using a single reference sample during inference. This work compares these two adaptation approaches, along with a training-from-scratch baseline, across different metrics. We show that fine-tuning a non-autoregressive architecture, such as ForwardTacotron, achieves speaker-adaptation quality comparable to zero-shot methods, with lower computational complexity but a lower mean opinion score. As a second contribution, we show that fine-tuning can be made more efficient in terms of data and while largely retaining quality.
Authors
- Christian Dittmar (ORCID: https://orcid.org/0000-0002-3220-2446)
- Frank Zalkow (ORCID: https://orcid.org/0000-0003-1383-4541)
- Emanuël A. P. Habets (ORCID: https://orcid.org/0000-0002-2613-8046)
- Nicola Pia (ORCID: https://orcid.org/0000-0003-0987-863X)
- Kishor Kayyar Lakshminarayana (ORCID: https://orcid.org/0000-0001-7493-818X)
Institutions
- International Audio Laboratories Erlangen (DE)
- Fraunhofer Institute for Integrated Circuits (DE)
Publication Details
- Journal
- EURASIP Journal on Audio Speech and Music Processing
- Published
- 2026-09-24
- DOI
- https://doi.org/10.1186/s13636-026-00479-w
- Primary Topic
- Speech Recognition and Synthesis
- Type
- article
- Field-Weighted Citation Impact
- 0.00