
Chung Tran, Sakriani Sakti · EURASIP Journal on Audio Speech and Music Processing 2026 · 2026
DOI: 10.1186/s13636-026-00474-1
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Recent advances have made it possible to generate speech directly from images without relying on text as a bridge, which is particularly beneficial for unwritten languages. These methods leverage speech units as intermediate representations and jointly train image-to-unit (I2U) and unit-to-speech (U2S) models, where the image encoder plays a critical role in visual semantic interpretation. However, existing encoders often fail to focus on salient image regions, leading to sub-optimal descriptions. In contrast, humans naturally employ a selective attention mechanism to focus on salient objects and their spatial relationships within a broader context, which produces precise and coherent descriptions. Inspired by this natural ability, we propose visual-saliency-to-speak (ViS2Speak), a novel framework that incorporates explicit visual and saliency information through the proposed visual-saliency guided multi-head attention (VSG-MHA) module. In addition, we introduce auxiliary acoustic loss (AuxLoss), which leverages ground-truth speech to optimize the training of the I2U model and enhance the overall efficiency of the image-to-speech (I2S) task. Experimental results demonstrate that this dual-support mechanism, combining visual saliency guidance and auxiliary acoustic loss, achieves competitive performance on I2S benchmarks such as Flickr8k, SpokenCOCO, and STAIR. Furthermore, the proposed framework also exhibits strong generalizability across diverse languages, including English, Japanese, Vietnamese, and Yoruba, under diverse data resource conditions.
No comments yet — start the discussion below.