Liuwen Li, Tao Wang, Yan Zhang, Nuan Zhang · Scientific Reports 2026 · 2026
DOI: 10.1038/s41598-026-70081-9
Scientific ReportsJournal465 h-indexCounts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
High-resolution remote sensing images (HRSIs) contain rich spatial details and fine-grained semantic information. However, inherent characteristics—including multi-scale object appearances, a high density of small objects, and low inter-class visual discriminability—significantly hinder the performance of general-domain Vision-Language Models (VLMs) in fine-grained remote sensing interpretation. To address these challenges, we propose RSFG-MoE , a mixture-of-experts vision-language model specifically designed for fine-grained interpretation of HRSIs. Building upon a general-domain Mixture-of-Experts (MoE) architecture, we introduce an Adaptive Semantic-Aware Tiling (ASAT) strategy to allocate visual computation to semantically informative regions while preserving fine-grained local features. We further introduce a dedicated fine-grained vision encoder , FG-CLIP, and a hierarchical feature fusion module to jointly capture global semantics and local structural details. To enhance computational efficiency, we integrate a DeepSeekMoE sparse language model; additionally, we adopt a two-stage progressive fine-tuning paradigm to adapt general-domain visual-language knowledge to remote sensing scenarios. Experiments on CHOICE, RSVQA-HR, RSICD, FAIR1M, NWPU-VHR-10, and MAR20 show that RSFG-MoE achieves task-dependent advantages in several fine-grained perception tasks, including image-level classification, scene-instance identification, attribute recognition, comparison-based visual question answering, and mAcc@0.5-based fine-grained recognition and localization. The results support the effectiveness of the proposed architecture for selected fine-grained remote-sensing interpretation tasks.
No comments yet — start the discussion below.