Mingchun Li, Xinran Wu, Rui Wang, Jining Bao · Electronics 2026 · 2026
DOI: 10.3390/electronics15184316
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Object detection based on visible-light images faces significant challenges under complex illumination and environmental conditions. Introducing infrared or depth images as complementary modalities can effectively enhance detection performance in such scenarios. However, most existing methods only support bi-modal configurations with identical branches for different modalities. Moreover, they apply generic fusion operations without adapting to the distinct characteristics of each modality. This limits their ability to fully exploit the complementary nature of RGB, IR, and Depth. To address these issues, we propose a multi-modal cross-attention fusion network named MCF-Net, based on cross-modal attention and difference convolution. The framework builds an asymmetric backbone that assigns lightweight branches to infrared and depth modalities. By introducing difference convolution, it explicitly extracts edge and structural features to compensate for the lack of texture information in these modalities. Building on this, a grid-based cross-attention mechanism is proposed via adaptive pooling within grid windows. It enables fine-grained cross-modal feature interaction while effectively controlling computational cost. Finally, a more refined spatial fusion module is developed to extract multi-scale spatial weights for each modality and adaptively fuse them via spatial attention. Our method is designed for tri-modal fusion and remains effective for bi-modal settings. Experiments on the AIC2026 challenge and the M3FD dataset show that our method achieves mAP50-95 of 39.0% and 59.7%, respectively. It achieves competitive results against existing methods while using significantly fewer parameters. These results validate the effectiveness of our multi-modal fusion strategy based on the proposed architecture.
No comments yet — start the discussion below.