Yijie Zhang, Jian Cheng, Shimiao Fan, Changjian Deng, Ziying Xia, Nyima Tashi · International Journal of Applied Earth Observation and Geoinformation 2026 · 2026
DOI: 10.1016/j.jag.2026.105439
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
The task of referring remote sensing image segmentation (RRSIS) aims to predict segmentation masks of target objectsbased on text descriptions. Achieving effective alignment between text modality and visual modality is crucial forimproving the performance of this task. However, most existing methods overlook the inherent domain differencesbetween the two modalities and directly fuse visual features with text features. Moreover, they lack the guidance ofdomain prior during the alignment process, which easily leads to inaccurate segmentation results. In this paper, wepropose a novel method for RRSIS—TIANet. This method combines implicit and explicit alignment approaches toachieve fine-grained image-text alignment. Moreover, to compensate for the lack of domain priors during the alignmentprocess, we leverage the Remote Sensing Vision Foundation Model (RS-VFM) to guide both alignment strategies. Onone hand, the Prior-Guided Visual-Text Feature Alignment Module (PVTAM) leverages domain priors to fuse textualand visual features in a stage-wise manner, thereby achieving cross-modal implicit alignment at the feature level. Onthe other hand, during training, explicit alignment between modalities is realized through the Dynamic BidirectionalAlignment and Prior Distillation (DBAPD) Loss. Furthermore, to enhance the perception ability of small targets,the Rotation-Adaptive Small Target Enhancement Module (RASTEM) is used in the decoder stage. We evaluate theproposed method on two public RRSIS datasets, and it achieves excellent performance compared with other advancedmethods.
No comments yet — start the discussion below.