Jiaxiang Fang, Shiqiang Ma, Jing Wang, Shengfeng He, Fei Guo · Neural Networks 2026 · 2026
DOI: 10.1016/j.neunet.2026.109694
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Vision-language models show promise in zero-shot semantic segmentation, but a key challenge is the disconnect between text and visual features. While text embeddings can roughly localize unseen objects, they often lack the fine-grained detail necessary for accurate segmentation, leading to oversegmentation or undersegmentation. To address this, we propose MODdapter, a Multimodal Adapter with Proportional-Integral-Derivative (PID) control for precise dual-space alignment of vision and language features. Our core idea is to leverage visual cues obtained during testing to complement textual information, improving the alignment between text and visual features. MODdapter extracts visual semantic cues from coarse segmentation results and integrates them as supplementary textual data, allowing for a more accurate projection of text features into the feature space and enhancing fine-grained recognition. To maintain stability and counteract potential oscillations caused by deviations in the extracted instance cues, we incorporate a PID control module that regulates the alignment process. PID control module leverages the combined joint constraints of short-term (derivative) and long-term (integral) memory, achieving fine-grained recognition of unseen target. This strategy significantly enhances the recognition of unseen objects, achieving state-of-the-art performance across multiple datasets. On the challenging COCO-20 i dataset, our method outperforms the current state-of-the-art by a significant improvements of 6.9% .
No comments yet — start the discussion below.