Huachen Lin, Zhiwei Fu, Xiumei Chen, Guirong Feng · Remote Sensing 2026 · 2026
DOI: 10.3390/rs18183175
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Visible-infrared (RGB-IR) object detection leverages multimodal information to ensure reliable perception in complex environments. However, dynamic scenes pose significant challenges due to the frequent inconsistency between scene-level modality contributions and local spatial reliability. Furthermore, standard feature extraction progressively attenuates boundary-sensitive structural cues, and unified fusion strategies often fail to capture spatially varying cross-modal complementarity. To overcome these limitations, we propose a Boundary-Guided Dual-Perspective Cross-Modal Fusion Network (BDPNet) to explicitly preserve shallow geometric structures and decouple deep semantic fusion into macroscopic and microscopic perspectives. Specifically, a Geometric Boundary Enhancement Module (GBEM) embeds Sobel-based high-frequency priors into shallow dual-modal features via residual spatial modulation, preventing the loss of crucial localization cues during downsampling. In the deep semantic space, a Hybrid Dual-Perspective Adaptive Fusion Module (HDAM) employs an illumination-aware branch for global modality weighting and a spatial confidence-driven branch for local cross-modal rectification. A spatial gating mechanism then dynamically reconciles these macro-environmental and micro-signal features. Extensive experiments on M3FD, LLVIP, and DroneVehicle demonstrate the effectiveness of BDPNet. Compared with state-of-the-art methods, BDPNet improves mAP50-95 by 0.8% and 1.0% on M3FD and LLVIP, respectively, and improves mAP50 by 0.6% on DroneVehicle, while using substantially fewer parameters and lower computational cost.
No comments yet — start the discussion below.