Minghong Xie, Hao Li, Yafei Zhang, Huafeng Li, Neng Dong, Yunbin Tu · Image and Vision Computing 2026 · 2026
DOI: 10.1016/j.imavis.2026.106155
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Remote Sensing Image Change Captioning (RSICC) aims to characterize the change information present in multi-temporal remote sensing images through natural language descriptions. However, existing methods have often struggled with accurate change localization and are susceptible to pseudo-changes caused by seasonal variations or illumination differences, primarily due to the absence of high-level logical guidance. To address this problem, we propose a Prior-Guided Evidence Decoupling Framework, which establishes a coarse-to-fine reasoning pathway. Specifically, we first design a Cognitive Prior Prompt Module (CPPM) to generate a macroscopic Cognitive Prior Prompt (CPP). This prompt serves as a valid logical premise, explicitly indicating the global change state to prevent hallucinations. Guided by this prior, we introduce a Difference Evidence Decoupling Module (DEDM) with a prior-guided decoupling loss.The DEDM disentangles bi-temporal visual evidence into common evidence (stable background) and private evidence (unique changes), thereby effectively isolating semantic changes from environmental noise. Finally, the Difference Evidence Fusion Module (DEFM) synthesizes these structured cues to drive a pre-trained language model for accurate caption generation. Extensive experiments on benchmark datasets demonstrate that our method significantly outperforms state-of-the-art approaches in both quantitative metrics and qualitative reasoning capabilities. The code is available at: https://github.com/LH2002lihao/repro-code .
No comments yet — start the discussion below.