
Yan Zhu, yu wang, Yijin Zhou, Bomin Liu, Rui Zhou · PLoS ONE 2026 · 2026
DOI: 10.1371/journal.pone.0358777
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Deep learning-based zero-shot speech synthesis has achieved substantial progress in speaker generalization, but stable modeling remains challenging in fine-grained emotional scenarios. Existing systems often process textual and emotional conditions through shared or closely coupled pathways, which may introduce interference between semantic content and emotional expression. In addition, uniformly averaging per-sample Conditional Flow Matching (CFM) losses may provide insufficient optimization emphasis to high-loss emotional samples. This study proposes a CFM-based zero-shot emotional speech synthesis method. Emotion–Text Decoupling Attention (ETDA) processes semantic and emotional conditions through parallel cross-attention streams and combines them through adaptive gated fusion, allowing the two conditions to retain their respective information before fusion. A 16-class fine-grained emotion space is constructed through classifier filtering and K-Means clustering based on pitch and energy features, and the resulting labels are mapped to continuous representations using a trainable lookup embedding. During training, Adaptive Loss-Threshold Reweighting estimates a threshold from mini-batch loss statistics and assigns larger weights to samples whose individual CFM losses exceed that threshold. Under speaker-disjoint evaluation on ESD, the proposed method achieves a word error rate of 3.62 ± 0.06 % , a mel-cepstral distortion of 4.45 ± 0.05 dB, and an emotional expressiveness mean opinion score of 4.64 ± 0.04 . Controlled text-length expansion and analyses on a fixed subset of high-loss emotional samples further indicate that the proposed method maintains linguistic content and emotional expression more consistently under the evaluated conditions. These results suggest that the proposed method can, to some extent, improve content preservation, emotional expression, and acoustic reconstruction in fine-grained zero-shot emotional speech synthesis.
No comments yet — start the discussion below.