Yulin Cheng, Haizheng Yu, Hong Bian · Image and Vision Computing 2026 · 2026
DOI: 10.1016/j.imavis.2026.106210
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Multispectral object detection exploits the complementary texture and thermal information in RGB and infrared images. Effective fusion, however, remains challenging when spatial correspondence is imperfect or RGB features deteriorate under changing illumination. We propose GRIFNet, a framework that couples prior-guided graph interaction with adaptive modality fusion. The Illumination-Guided Graph Interaction Module (IG-GIM) extracts learned luminance-related and gradient-based structural representations from raw RGB patches without illumination annotations. These priors guide the selection of non-local graph neighborhoods, extending their role beyond feature weighting or concatenation and reducing reliance on semantic-feature similarity alone. Graph message passing then supplies contextual information beyond rigid same-location correspondence, without explicit geometric registration. The subsequent Structure-Enhanced Gated Fusion (SEGF) module combines high-frequency residual injection and pixel-wise competitive gating to enhance local structural detail and adapt the contributions of RGB and infrared features. Experiments on M3FD, FLIR, LLVIP, and VEDAI demonstrate competitive detection performance in ground-level and aerial-view settings. Controlled ablations distinguish prior feature concatenation from topology guidance, showing a larger gain from prior-guided neighborhood selection and a complementary benefit from SEGF. Cross-modal perturbation tests, capacity-controlled comparisons, and inference measurements further assess detection performance under spatial discrepancies and the computational trade-offs of the framework.
No comments yet — start the discussion below.