A pronunciation-sensitive analysis of XTTS adaptation for modern standard Arabic synthetic diacritization length rescue and training dynamics

Abstract Modern Standard Arabic (MSA) text-to-speech evaluation is complicated by the omission of short vowels and case endings in ordinary writing: a system may remain lexically intelligible while realizing a linguistically inappropriate pronunciation. This study examines XTTS v2 adaptation through matched raw, fully synthetic-diacritized, partially diacritized, and punctuation-guided length-rescue conditions. The retained Arabic speech pool contains 99,140 audio segments (296.94 h, 24 kHz mono), while the controlled adaptations use fixed subsets of 0.805–4.245 h. Evaluation combines a 60-sentence diagnostic benchmark, canonical ASR-proxy WER/CER, 99 targeted case-ending judgments, blinded expert naturalness and pronunciation ratings, pairwise preferences, and checkpoint analysis at 1, 3, and 13 total epochs. In the canonical evaluation, the off-the-shelf model was the strongest benchmark-wide reference (WER 0.224; CER 0.071), and every controlled adapted condition had a higher modeled overall error rate after Holm correction. Rescue-M produced the lowest descriptive case-sensitive CER (0.065), narrowly below B0 (0.067), but none of its planned case-sensitive contrasts remained statistically distinguishable after correction. An unseen explicit token substantially increased WER and CER (rate ratios 2.159 and 3.515). Validation loss declined across the observed checkpoints, whereas external error trajectories were non-monotonic and configuration-dependent. The independent evaluator produced a different case-ending ordering from R1, while pronunciation ratings favored B0 more consistently than naturalness or case-ending scores. The findings therefore support a diagnostic rather than leaderboard interpretation: the apparent benefit of Arabic XTTS adaptation depends materially on the evaluation metric, checkpoint, and evaluator.

Authors

Institutions

Publication Details

Journal
Discover Artificial Intelligence
Published
2026-09-24
DOI
https://doi.org/10.1007/s44163-026-02214-y
Primary Topic
Speech Recognition and Synthesis
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

A pronunciation-sensitive analysis of XTTS adaptation for modern standard Arabic synthetic diacritization length rescue and training dynamics

Hamid Tairi, Abdennabi Morchid, Mohamed-Amine Chadi, Khalid Oqaidi et al.
Discover Artificial Intelligence
Speech Recognition and Synthesis
article

A pronunciation-sensitive analysis of XTTS adaptation for modern standard Arabic synthetic diacritization length rescue and training dynamics

Hamid Tairi, Abdennabi Morchid, Mohamed-Amine Chadi, Khalid Oqaidi, Zouhair Elamrani Abou Elassad, Abdennacer Elbasri, Ali Yahyaouy
article en

Abstract

Abstract Modern Standard Arabic (MSA) text-to-speech evaluation is complicated by the omission of short vowels and case endings in ordinary writing: a system may remain lexically intelligible while realizing a linguistically inappropriate pronunciation. This study examines XTTS v2 adaptation through matched raw, fully synthetic-diacritized, partially diacritized, and punctuation-guided length-rescue conditions. The retained Arabic speech pool contains 99,140 audio segments (296.94 h, 24 kHz mono), while the controlled adaptations use fixed subsets of 0.805–4.245 h. Evaluation combines a 60-sentence diagnostic benchmark, canonical ASR-proxy WER/CER, 99 targeted case-ending judgments, blinded expert naturalness and pronunciation ratings, pairwise preferences, and checkpoint analysis at 1, 3, and 13 total epochs. In the canonical evaluation, the off-the-shelf model was the strongest benchmark-wide reference (WER 0.224; CER 0.071), and every controlled adapted condition had a higher modeled overall error rate after Holm correction. Rescue-M produced the lowest descriptive case-sensitive CER (0.065), narrowly below B0 (0.067), but none of its planned case-sensitive contrasts remained statistically distinguishable after correction. An unseen explicit token substantially increased WER and CER (rate ratios 2.159 and 3.515). Validation loss declined across the observed checkpoints, whereas external error trajectories were non-monotonic and configuration-dependent. The independent evaluator produced a different case-ending ordering from R1, while pronunciation ratings favored B0 more consistently than naturalness or case-ending scores. The findings therefore support a diagnostic rather than leaderboard interpretation: the apparent benefit of Arabic XTTS adaptation depends materially on the evaluation metric, checkpoint, and evaluator.

Discover Artificial IntelligenceVol. 6(1)
Cadi Ayyad University (MA), Daffodil International University (BD), Université Hassan II Mohammedia (MA), Sidi Mohamed Ben Abdellah University (MA), University of Hassan II Casablanca (MA)
Quality Education
Openalex Percentile: Top 9%
Speech Recognition and Synthesis
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.