XueZhuan Zhao, Zhenhao Zhao, Lingling Li, Xiaoyan Shao, MengMeng Tang, Xiaoming Bai · Complex & Intelligent Systems 2026 · 2026
DOI: 10.1007/s40747-026-02474-2
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Weakly supervised semantic segmentation (WSSS) aims to produce pixel-level predictions from image-level labels, but it often suffers from incomplete object activation and background ambiguity due to the inherent bias of class activation maps (CAMs). Existing CLIP-based methods improve semantic alignment, they struggle to jointly capture fine grained local details and long-range global dependencies, leading to fragmented activations and blurred boundaries. To address this, we propose a CLIP-based WSSS framework with multi-scale semantic enhancement attention (MSEA) module. MSEA combines multi-scale depthwise separable convolution (DSConv) for local feature extraction, linear attention for efficient global interaction, and an efficient attention (EA) refinement mechanism to suppress noise and enhance boundary quality. And adopt structured attribute embeddings for better semantic guidance. Experiments show that our method outperforms existing single-stage approaches, achieving 75.4%, 76.6%, and 48.1% mIoU on PASCAL VOC 2012 and MS COCO 2014, with clear improvements in CAM completeness and segmentation quality.
No comments yet — start the discussion below.