Dongdong Bu, Lei Liu, Zhuli Xie, Gang Wan · Preprints.org 2026 · 2026
DOI: 10.20944/preprints202609.1269.v1
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Remote sensing image captioning enables automated interpretation of remote sensing data for applications such as environmental monitoring and urban planning. However, general-purpose multimodal models exhibit suboptimal adaptability on remote sensing imagery due to domain discrepancies in data distribution, resolution, and rotation invariance. We propose a domain-specific memory bank-enhanced framework for remote sensing image caption, built upon the lightweight EVCap architecture with three core innovations: (1) a remote sensing specific visual memory bank providing domain-representative visual-semantic alignment; (2) a staged transfer training strategy that leverages pre-trained knowledge while mitigating overfitting from limited remote sensing data; and (3) a K-means clustering-based few-shot memory bank that preserves performance while reducing memory footprint for resource-constrained deployment. Extensive experiments on NWPU, RSICD, and Sydney datasets demonstrate substantial improvements over baselines. Specifically, CIDEr increases by 3.3% on RSICD (0.2735 to 0.2825) and 11.7% on Sydney (0.1422 to 0.1589). The proposed method also achieves competitive performance against state-of-the-art methods while maintaining favorable computational efficiency.
No comments yet — start the discussion below.