
Abdennacer Elbasri, Abdennabi Morchid, Khalid Oqaidi, Hamid Tairi, Ali Yahyaouy, Mohamed-Amine Chadi, Zouhair Elamrani Abou Elassad · Discover Artificial Intelligence 2026 · 2026
DOI: 10.1007/s44163-026-02214-y
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Modern Standard Arabic (MSA) text-to-speech evaluation is complicated by the omission of short vowels and case endings in ordinary writing: a system may remain lexically intelligible while realizing a linguistically inappropriate pronunciation. This study examines XTTS v2 adaptation through matched raw, fully synthetic-diacritized, partially diacritized, and punctuation-guided length-rescue conditions. The retained Arabic speech pool contains 99,140 audio segments (296.94 h, 24 kHz mono), while the controlled adaptations use fixed subsets of 0.805–4.245 h. Evaluation combines a 60-sentence diagnostic benchmark, canonical ASR-proxy WER/CER, 99 targeted case-ending judgments, blinded expert naturalness and pronunciation ratings, pairwise preferences, and checkpoint analysis at 1, 3, and 13 total epochs. In the canonical evaluation, the off-the-shelf model was the strongest benchmark-wide reference (WER 0.224; CER 0.071), and every controlled adapted condition had a higher modeled overall error rate after Holm correction. Rescue-M produced the lowest descriptive case-sensitive CER (0.065), narrowly below B0 (0.067), but none of its planned case-sensitive contrasts remained statistically distinguishable after correction. An unseen explicit token substantially increased WER and CER (rate ratios 2.159 and 3.515). Validation loss declined across the observed checkpoints, whereas external error trajectories were non-monotonic and configuration-dependent. The independent evaluator produced a different case-ending ordering from R1, while pronunciation ratings favored B0 more consistently than naturalness or case-ending scores. The findings therefore support a diagnostic rather than leaderboard interpretation: the apparent benefit of Arabic XTTS adaptation depends materially on the evaluation metric, checkpoint, and evaluator.
No comments yet — start the discussion below.