Dongsuk Yook, Semin Kim, Hyung-Pil Chang · Applied Sciences 2026 · 2026
DOI: 10.3390/app16189314
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
This paper proposes a novel hybrid voice conversion framework, termed VQ-CycleDiffusion, which integrates vector quantization (VQ) with CycleDiffusion to address the inherent limitations of continuous latent representations in diffusion-based models. While conventional diffusion-based voice conversion (VC) models provide superior speech quality, they often suffer from speaker information leakage due to their reliance on continuous latent spaces, which fail to completely separate linguistic content from source speaker characteristics. To overcome this issue, we introduce a discrete bottleneck via VQ to extract multi-speaker linguistic representations, which are then used as conditional inputs for CycleDiffusion after a carefully designed codeword transformation between the source and target speakers. Experimental results on the VCTK corpus demonstrate that the proposed model significantly outperforms the conventional method in terms of speaker similarity, as confirmed by both i-vector and x-vector cosine similarity metrics, while maintaining spectral reconstruction performance comparable to that of the baseline model, as evidenced by stable Mel-cepstral distance (MCD) values. Furthermore, the quality of the converted speech was not degraded, as demonstrated by the predicted mean opinion score (MOS) evaluation. Ablation studies reveal that initializing the codebook using k-means clustering and focusing on diffusion model fine-tuning play key roles in maximizing performance. These results validate that the proposed method effectively disentangles linguistic content from speaker characteristics while preserving the high-fidelity generation capability of diffusion models.
No comments yet — start the discussion below.