Dong Xing, Hang Yang, Jinhe Zhang, Peixun Liu, Yuqing Wang · Cognitive Computation 2026 · 2026
DOI: 10.1007/s12559-026-10661-z
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
High-resolution remote sensing semantic segmentation is a key enabler for fine-grained urban mapping and related geospatial applications. However, complex textures, severe scale variations, and cross-modal discrepancies make models with fixed architectures and computation paths struggle to combine reliable multimodal fusion with adequate global-context modeling. To address these challenges, we propose RL-Axial, a multimodal segmentation framework with interpretable, sample-adaptive context aggregation. A cross-modal gated fusion head (CMGF) performs pixel-wise fusion between RGB and digital surface model (DSM) features across multiple scales. A controllable axial-attention operator then provides four alternatives— off , row , col , and row+col —instead of applying a single axial configuration to every image patch. We train a lightweight selector with a group-normalized policy-gradient objective based on per-sample relative rewards. During the second training stage, the modality encoders and CMGF are frozen, whereas the policy, axial operator, decoder, and segmentation head are updated through their respective policy and segmentation losses. On the evaluated ISPRS Vaihingen and Potsdam splits, a single run of RL-Axial obtains 84.58 and 86.73 mIoU, respectively, numerically 0.35 and 0.53 points above the literature-reported FTransUNet values. The source-specific baseline protocols preclude a controlled superiority claim. Qualitative examples also show cleaner boundaries and more coherent object structures.
No comments yet — start the discussion below.