GUO Junjun ZHANG Kaili · DOAJ (DOAJ: Directory of Open Access Journals) 2026 · 2026
DOI: 10.3778/j.issn.1673-9418.2509039
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Domain-specific multimodal neural machine translation (DMNMT) incorporates images as an additional modality input to achieve accurate translation from the source language to the target language within specific domains, thereby enhancing translation accuracy and robustness. Since visual information contains rich domain-specific details, its integration can strengthen the domain representation capabilities of model and further optimize the translation of domain-specific terminology. Most existing MNMT approaches enhance text with visual information through fine-grained cross-modal architectures, but they fail to selectively mine and filter domain-related and non-domain information. This results in insufficient capability of the image-text fusion architecture to screen and integrate domain visual features, leading to target translations lack of strong domain relevance. To address the challenge of filtering and integrating visual information at different domain levels, this paper proposes a text-image fusion method based on visually adaptive segmentation and contrastive enhancement. This method progressively disentangles domain-related visual information from non-domain visual information associated with the text, improving the capture of domain information. Additionally, it introduces a source language-domain visual-non-domain visual cross-modal and cross-lingual triplet constraint that pulls domain information closer to the source language while pushing non-domain information farther away. This effectively boosts the performance of domain-specific multimodal machine translation. Experimental results on the Fashion-MMT and EMMT domain datasets, as well as the general-domain Multi30K dataset, demonstrate the effectiveness of the proposed method. Further ablation studies and visualizations confirm the efficacy and robustness of model.
No comments yet — start the discussion below.