A pronunciation-sensitive analysis of XTTS adaptation for modern standard Arabic synthetic diacritization length rescue and training dynamics
Abstract Modern Standard Arabic (MSA) text-to-speech evaluation is complicated by the omission of short vowels and case endings in ordinary writing: a system may remain lexically intelligible while realizing a linguistically inappropriate pronunciation. This study examines XTTS v2 adaptation through matched raw, fully synthetic-diacritized, partially diacritized, and punctuation-guided length-rescue conditions. The retained Arabic speech pool contains 99,140 audio segments (296.94 h, 24 kHz mono), while the controlled adaptations use fixed subsets of 0.805–4.245 h. Evaluation combines a 60-sentence diagnostic benchmark, canonical ASR-proxy WER/CER, 99 targeted case-ending judgments, blinded expert naturalness and pronunciation ratings, pairwise preferences, and checkpoint analysis at 1, 3, and 13 total epochs. In the canonical evaluation, the off-the-shelf model was the strongest benchmark-wide reference (WER 0.224; CER 0.071), and every controlled adapted condition had a higher modeled overall error rate after Holm correction. Rescue-M produced the lowest descriptive case-sensitive CER (0.065), narrowly below B0 (0.067), but none of its planned case-sensitive contrasts remained statistically distinguishable after correction. An unseen explicit token substantially increased WER and CER (rate ratios 2.159 and 3.515). Validation loss declined across the observed checkpoints, whereas external error trajectories were non-monotonic and configuration-dependent. The independent evaluator produced a different case-ending ordering from R1, while pronunciation ratings favored B0 more consistently than naturalness or case-ending scores. The findings therefore support a diagnostic rather than leaderboard interpretation: the apparent benefit of Arabic XTTS adaptation depends materially on the evaluation metric, checkpoint, and evaluator.
Authors
- Hamid Tairi (ORCID: https://orcid.org/0000-0002-5445-0037)
- Abdennabi Morchid (ORCID: https://orcid.org/0000-0001-9349-2822)
- Mohamed-Amine Chadi (ORCID: https://orcid.org/0000-0003-3260-0978)
- Khalid Oqaidi (ORCID: https://orcid.org/0000-0002-7223-3298)
- Zouhair Elamrani Abou Elassad (ORCID: https://orcid.org/0000-0003-0051-9374)
- Abdennacer Elbasri (ORCID: https://orcid.org/0009-0002-0881-7974)
- Ali Yahyaouy
Institutions
- Cadi Ayyad University (MA)
- Daffodil International University (BD)
- Université Hassan II Mohammedia (MA)
- Sidi Mohamed Ben Abdellah University (MA)
- University of Hassan II Casablanca (MA)
Publication Details
- Journal
- Discover Artificial Intelligence
- Published
- 2026-09-24
- DOI
- https://doi.org/10.1007/s44163-026-02214-y
- Primary Topic
- Speech Recognition and Synthesis
- Type
- article
- Field-Weighted Citation Impact
- 0.00