Hien Tran-Hy Luong, Văn Thế Thành · International journal of intelligent engineering and systems 2026 · 2026
DOI: 10.22266/ijies2026.1031.36
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Knowledge-enhanced semantic image retrieval aims to handle relational queries that remain challenging for vision-language models.However, many reported improvements are evaluated using ground-truth scene graphs, making their practical value unclear.This paper adapt the Open Knowledge Format, a standardized representation that separates visual extraction from graph-based reasoning and enables controlled analysis.Experiments on Visual Genome and MS-COCO evaluate four extraction conditions from ground-truth annotations to fully predicted objects and relations.Oracle scene graphs reach 41.3 percent mean average precision at ten, 8.2 times the fully predicted pipeline and 7.7 times the CLIP baseline.A controlled sweep shows that the contribution of relational evidence decays from 0.70 to 0.01 points as entity recall falls from 100 to 20 percent, identifying visual extraction quality as the primary bottleneck.This is a diagnostic study of a single controlled protocol; no competitive advantage is claimed.
No comments yet — start the discussion below.