Hongru Xiao, Bo Li, Bin Yang, Jinming Hu, Jianquan Li, Yan Kai, Hong Li, Xiang Li · Measurement Science and Technology 2026 · 2026
DOI: 10.1088/1361-6501/ae9207
Measurement Science and TechnologyJournal182 h-indexCounts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
RGB-infrared (RGB-IR) fusion improves object detection by enabling robust object localization under challenging illumination and environmental conditions. However, the additional IR modality often increases computational cost, limiting its deployment in real-time measurement and perception systems. This work revisits cross-modal fusion strategies and shows that strong performance can be achieved by progressively accumulating efficient cross-modal interactions across multiple feature scales, without duplicating auxiliary modality feature-extraction branches. To exploit this observation, a cross-modal fusion (CMFusion) detection framework is proposed, which focuses on efficient multi-level RGB-IR feature interaction within YOLO-based detectors. This design reduces redundant feature extraction within individual modalities while enhancing multi-level multimodal interaction, enabling effective cross-modal complementarity with lower computational overhead. Extensive experiments on three public RGB-IR datasets across two detection tasks (OBB and HBB) demonstrate that CMFusion achieves a 1.7%–6.1% mAP improvement over baseline unimodal detectors while maintaining comparable low computational complexity, and consistently surpasses state-of-the-art methods.
No comments yet — start the discussion below.