A. Bala Raju, Sp Singh, Dhiraj Sunehra · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22842363
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Recent advancements in voice conversion systems requiring parallel data. These models have demonstrated have been largely driven by deep learning techniques, enabling notable success in preserving speaker identity and linguistic the high-quality synthesis of human speech. However, existing content while enabling flexible and speaker-independent models often fail to generate emotionally expressive speech, conversions. The main goal of this research is to fill this gap by incorporating a prosody-aware extension into embedding- guided neural voice conversion (EGNVC). This extension seeks to improve the emotional expressiveness and naturalness of the converted speech, while preserving both speaker identity and linguistic content. A new framework is introduced that integrates a content encoder, speaker embedder, and prosody extractor to address this challenge. The prosody extractor is designed to capture dynamic prosodic elements like pitch, energy, and timing from a reference audio signal. These features are then introduced into the decoding process through a prosody conditioning module, allowing for precise control over prosody during speech synthesis. FiLM (Feature-wise Linear Modulation) layers are employed to adjust intermediate features based on speaker and prosody embeddings, ensuring that both the speaker's identity and the intended prosodic qualities are accurately maintained. The vocoder converts the generated mel-spectrograms into high- quality waveforms, ensuring a natural-sounding output. Experimental results show that the proposed model surpasses conventional voice conversion systems in maintaining linguistic content and prosodic features, with significant improvements in Mel-Cepstral Distortion (MCD), Pitch RMSE, and Energy RMSE. Subjective assessments, including Mean Opinion Score (MOS) and Emotion Classification Accuracy, highlight the model's enhanced emotional expressiveness and naturalness. These findings emphasize the critical role of prosody in high- quality voice conversion, offering a robust framework for creating more dynamic, expressive, and human-like speech synthesis systems.
No comments yet — start the discussion below.