Xin Zhang, Yu Liu, Shimin Shan, Zhehuan Zhao · Advanced Engineering Informatics 2026 · 2026
DOI: 10.1016/j.aei.2026.105137
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Multi-modal entity alignment aims to identify equivalent entities across two different multi-modal knowledge graphs. Most traditional multi-modal entity alignment methods rely on access to entity names or relation names as a prerequisite, making them unsuitable for industrial domains where these surface forms are unavailable, such as the aerospace and military industries, where entity names are often withheld for confidentiality. This study proposes a cross-modal interactive entity alignment framework designed specifically for surface-form-unavailable settings. The framework first generates entity and relation embeddings based on structural features. Then, a cross-modal interactive strategy, leveraging the attention mechanism, is designed to enhance both structural and visual features. Multi-modal entity alignment is achieved by integrating the enhanced structural and visual features through a weighted fusion strategy. In addition, the framework introduces a dynamic-guided iterative learning strategy during training, based on the similarity of individual feature embeddings, to further improve model performance. We evaluated the proposed CMIEA framework on five benchmark datasets, without utilizing entity names or relation names. Experimental results demonstrate that our framework significantly outperforms baseline methods, highlighting its potential and feasibility for applications in industrial domains under surface-form-unavailable settings.
No comments yet — start the discussion below.