Haocheng Li, Juepeng Zheng, Shuangxi Miao, Ruibo Lu, Guosheng Cai, Haohuan Fu, Jianxi Huang · International Journal of Applied Earth Observation and Geoinformation 2026 · 2026
DOI: 10.1016/j.jag.2026.105520
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Multimodal remote sensing semantic segmentation improves scene understanding by integrating complementary information from heterogeneous Earth observation data. Although pretrained Vision Foundation Models (VFMs) provide strong general-purpose representations, adapting them to multimodal tasks typically requires updating large modality-specific encoders, and optimization tends to be dominated by the optical modality, leaving auxiliary cues under-utilized. To address these challenges, we propose MoBaNet, a parameter-efficient multimodal adaptation framework that keeps the VFM backbone frozen and concentrates trainable capacity on three complementary components. First, the Cross-modal Prompt-Injected Adapter (CPIA) generates semantic prompts from paired modalities and injects them into lightweight bottleneck adapters to enable cross-modal interaction under the frozen backbone. Second, the Difference-Guided Gated Fusion Module (DGFM) exploits cross-modal discrepancy to guide channel-wise and spatially adaptive feature fusion. Third, the training-only Modality-Conditional Random Masking (MCRM) strategy locally corrupts one modality at a time and applies hard-pixel auxiliary supervision to modality-specific branches, encouraging more robust use of complementary cues. Extensive experiments on the ISPRS Vaihingen and Potsdam benchmarks show that MoBaNet achieves the best overall performance among the compared methods, while its default DINOv2-ViT-B configuration uses only 5.43% trainable parameters and remains robust under partial modality degradation.
No comments yet — start the discussion below.