ju jiang, Darko B. Vukovic, Jie Cao, Haoran Xie, Youquan Wang, Jia Wu · ACM Transactions on Multimedia Computing Communications and Applications 2026 · 2026
DOI: 10.1145/3802546
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Vision-Language Pre-training (VLP) aims to convert vision-language tasks into image-text matching challenges, necessitating robust interaction between visual and textual modalities. However, current research often faces a trade-off: methods pursuing fine-grained alignment typically rely on computationally expensive object detectors and pre-defined annotations, while coarse-grained alignment methods overlook subtle visual details. To alleviate these issues, we present OMEGA, a C O mprehensive Multi-l E vel G ranularity A lignment framework. OMEGA achieves sufficient cross-modal alignment without relying on external object detectors or annotations. Specifically, we first design a Text-Aware Patch Selection (TAPS) module, which dynamically constructs patch-level regions for fine-grained excavation based on informative token selection, effectively bypassing the need for bounding boxes. Secondly, to facilitate multi-grained information fusion, we introduce a Cross-Grained Alignment (CGA) module to learn modality-shared features across global and object levels. Experimental results demonstrate that OMEGA surpasses existing approaches in five pivotal vision-language tasks—image-text retrieval, visual question answering, visual reasoning, visual entailment, and image captioning, while maintaining superior inference efficiency compared to detector-based methods. We believe this study offers a promising object-free paradigm for scalable vision-language pre-training.
No comments yet — start the discussion below.