Yuliang Cai, Jesse Thomason, Mohammad Ghomi Rostami · · 2026
DOI: 10.1145/3776574.3831142
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Multimodal interactive systems increasingly rely on vision-language models (VLMs), such as CLIP, to mediate between users and visual content. These VLMs power image search and retrieval interfaces where users specify what they want in natural language. However, such systems suffer from a basic aspect of human communication: negation. When users say "a living room with a sofa but no TV," current vision-language models ignore the exclusion and return content that contradicts the user’s stated intent. This failure undermines user trust and forces costly correction loops in any interface that takes natural-language input. We address this gap from two directions. First, we introduce TNG-CLIP, a training-time negation data generation pipeline that equips CLIP with negation understanding while adding only 2.8% training time overhead, avoiding the prohibitive cost of LLM-generated negation datasets used in prior work. Second, we introduce NEG-T2I, the first large-scale benchmark for evaluating whether text-to-image generation correctly follows user-specified inclusions and exclusions. Across image-to-text matching, text-to-image retrieval, and constraint-aware generation, TNG-CLIP achieves SoTA results, narrowing the gap between user intent and system behavior in language-driven visual interfaces.
No comments yet — start the discussion below.