Hui Xiong, Wen Luo, Bing He · Journal of Imaging 2026 · 2026
DOI: 10.3390/jimaging12090453
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Remote sensing referring image segmentation faces critical challenges including arbitrary target rotation, drastic scale variation, cluttered complex backgrounds, and large visual-language semantic gaps. Existing mainstream segmentation models adopt fixed-receptive-field backbones, coarse unidirectional cross-modal interaction and static learnable object queries, which easily cause small-object missed detection, blurred boundary segmentation and fragmented predictions. This paper proposes a context-aware referring expression segmentation model for remote sensing to tackle the above limitations. We construct a dual-stream feature extraction backbone with InternImage and CLIP text encoder, and design a semantic prior localization map module to generate spatial heatmaps for spatial inductive bias and improve small-object localization recall. A cross-modal context aggregator performs multi-scale bidirectional visual-text alignment, while a dynamic query initialization strategy and a language-guided Transformer decoder progressively improve target localization and mask refinement. Experiments are conducted on two standard RRSIS benchmarks, RefSegRS and RRSIS-D. On RefSegRS, the proposed model achieves an mIoU of 70.43% and an oIoU of 77.34%, and obtains significant improvements on small vehicles, buildings and slender road markings. On the larger-scale RRSIS-D dataset, our model achieves an mIoU of 63.71% and an oIoU of 73.68%. The proposed model provides a solution for accurate multi-scale target segmentation guided by natural language descriptions in complex remote sensing scenes.
No comments yet — start the discussion below.