Lin Qi, Tiancun Guo, Yanfei Dong, Changbing Zheng, Mingliang Gao · Electronics 2026 · 2026
DOI: 10.3390/electronics15194440
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Single-modal object Re-identification (Re-ID) frequently suffers from performance degradation under complex and dynamic scene conditions. Although multi-modal object Re-ID leverages complementary information across diverse modalities to alleviate this issue, existing approaches remain highly susceptible to irrelevant background interference. To address these challenges, we propose the Spatial–Frequency Token Interaction Network (SFTINet) for object Re-ID. Specifically, SFTINet utilizes a Vision Transformer (ViT) to extract multi-modal tokens, which are subsequently processed by a Spatial–Frequency Token Interaction Module. This module captures fine-grained local details in the spatial domain while concurrently encoding global structural information in the frequency domain, thereby augmenting the multi-modal fine-grained representation capability. A dynamic token interaction mechanism further regulates cross-modal information flow via attention, facilitating the deep fusion of RGB, Near-Infrared (NIR), and Thermal Infrared (TIR) features. Additionally, a Reconstruction Loss Module minimizes the distribution discrepancy across modalities through token-level constraints, and a Token Mask Generation Module encourages foreground-centric learning by strategically masking key tokens. By seamlessly integrating spatial–frequency interaction with token-level masking, SFTINet effectively suppresses background interference while preserving cross-modal complementary information. Experimental results demonstrate that SFTINet achieves competitive performance on multi-modal benchmarks, validating its effectiveness in robust cross-modal feature fusion and fine-grained semantic modeling.
No comments yet — start the discussion below.