Kaiyu Li, Zixuan Jiang, Xiangyong Cao, Jiayu Wang, Yuchen Xiao, Jing Yao, Chen Wu, Deyu Meng, Zhi Wang · ISPRS Journal of Photogrammetry and Remote Sensing 2026 · 2026
DOI: 10.1016/j.isprsjprs.2026.09.019
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Automated textual description of remote sensing images is crucial for unlocking their full potential in diverse applications, from environmental monitoring to urban planning and disaster management. However, existing studies in remote sensing image captioning primarily focus on the image level, lacking object-level fine-grained interpretation, which prevents the full utilization and transformation of the rich semantic and structural information contained in remote sensing images. Consequently, we propose Geo-DLC, a novel task for fine-grained object-level image captioning in remote sensing. To systematically establish and evaluate this new paradigm, we introduce a comprehensive suite encompassing a dataset, a benchmark, and a specialized model. Specifically, we construct DE-Dataset, a large-scale collection comprising 25 categories and 261,806 annotated instances with detailed descriptions of object attributes, relationships, and contexts. To provide a rigorous evaluation standard, we develop DE-Benchmark, an LLM-assisted evaluation protocol based on question answering, to accurately measure the capabilities of models on the Geo-DLC task. Furthermore, we present DescribeEarth, an MLLM architecture explicitly tailored for this task. It integrates a scale-adaptive focal strategy and a domain-guided fusion module, which leverage the features of remote sensing Vision-Language Model (VLM) to encode high-resolution details and category priors while maintaining the global context. Extensive evaluations show that DescribeEarth achieves the highest aggregate scores on the simple and complex subsets and matches GPT-4o on the out-of-distribution subset of DE-Benchmark. Its clearest gains occur in surrounding-context coverage and in categories whose interpretation depends strongly on overhead spatial layout and remote-sensing context. All data, code, and weights are publicly available at https://github.com/earth-insights/DescribeEarth .
No comments yet — start the discussion below.