
Hongwei Liu, Yisha Liu, Weimin Xue, Yan Zhuang · Measurement Science and Technology 2026 · 2026
DOI: 10.1088/1361-6501/aea885
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Visible-thermal (RGB-T) imaging systems provide crucial complementary information for robust visual sensing and measurement in complex illumination conditions. However, effectively fusing multi-modal sensory data to achieve accurate semantic segmentation remains a key challenge. Most existing methods rely on heuristic fusion strategies or standard attention mechanisms, which fail to decouple modality-shared semantics from modality-specific details leading to redundant representations. Furthermore, conventional decoders progressively upsample features in a stage-wise manner, which may introduce cross-scale inconsistencies due to mismatched spatial resolutions and receptive fields. To address these challenges in multi-sensor data processing, we propose ODCANet, an Orthogonal Feature Decoupling and Hierarchical Cross-Scale Aggregation Network, which consists of an Orthogonal Feature Decoupling Module (OFDM) and a Hierarchical Cross-Scale Aggregation Module (HCSAM). Specifically, OFDM employs Correlation-guided Decomposition Blocks (CDB) with global-local correlation modeling to estimate shared and modality-specific components, and applies an explicit orthogonality constraint during training to encourage lower correlation between them, thereby promoting more complementary fused representations. Meanwhile, HCSAM performs dynamic cross-scale aggregation of adjacent-stage features to help alleviate semantic gaps across decoder stages. Extensive experiments on benchmark datasets demonstrate that ODCANet achieves the best reported mIoU among the compared methods, reaching 62.21%, 89.25% and 67.98% mIoU on MFNet, PST900 and FMB datasets, respectively.
No comments yet — start the discussion below.