Zhonghua Dang, Hui CHEN, Xianxun Zhu · Image and Vision Computing 2026 · 2026
DOI: 10.1016/j.imavis.2026.106185
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Unsupervised image clustering remains challenging because visually similar categories often lack explicit semantic boundaries, while purely visual representations are easily affected by appearance variations, background noise, and ambiguous local structures. Although recent image–text clustering methods introduce external textual semantics through vision–language models and lexical knowledge bases, their contrastive optimization still largely depends on static visual views and manually designed augmentations, which may provide insufficient or noisy positive samples. To address these limitations, we propose Diffusion-Augmented Adaptive Image–Text Contrastive Learning (DAITC), a unified unsupervised multimodal clustering framework that integrates adaptive text counterpart construction with diffusion-based semantic positive generation. Specifically, CLIP is first used to extract image and text embeddings, and a compact semantic vocabulary is constructed from WordNet according to image-level semantic centers. An adaptive temperature mechanism is then introduced to generate image-specific text counterparts by dynamically weighting candidate noun embeddings according to their similarity distributions. Beyond conventional visual augmentation, we further design a text-guided conditional diffusion module that generates semantically consistent positive views conditioned on both visual embeddings and adaptive text counterparts. These diffusion-augmented samples are incorporated into a reliability-aware soft contrastive objective, together with direct image–text alignment and neighborhood-level cross-modal consistency constraints. Finally, cluster balance and confidence regularization are employed to obtain stable and discriminative cluster assignments. Experiments are conducted on STL-10, CIFAR-10, CIFAR-20, DTD, and UCF-101 to evaluate clustering accuracy, semantic alignment, robustness, and generalization ability. The proposed framework provides a reliable way to combine external textual semantics and generative positive augmentation for unsupervised multimodal image clustering.
No comments yet — start the discussion below.