Xiyi Ban, Jiaying Yu, Fusheng Li, Bin Ai · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22803752
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
With the continuous development of Intelligent Transportation Systems (ITS), autonomous driving technologies have become increasingly prevalent in modern transportation. However, existing 3D object detection methods still suffer from depth information loss and misalignment between different modalities, leading to insufficient detection accuracy and robustness. To address these challenges, this paper proposes a novel multimodal 3D object detection model, termed FusionMultiNet. Specifically, FusionMultiNet integrates geometric information from LiDAR with semantic features from RGB images. By introducing the Multi-Depth Unprojection (MDU) strategy and the Gated Modality-Aware Convolution (GMA-Conv) module, the proposed model effectively resolves the issue of depth information loss while jointly alleviating feature fusion biases caused by modality misalignment. The MDU strategy achieves modality alignment in physical space through multi-depth unprojection, whereas the GMA-Conv module adaptively fuses image semantic features under the guidance of LiDAR geometric cues, significantly improving detection accuracy and robustness. In addition, FusionMultiNet incorporates a Modality-Specific Context Encoder (MSCE) to further enhance object feature representation. Experimental results demonstrate that FusionMultiNet consistently outperforms existing state-of-the-art methods on the nuScenes and KITTI benchmark datasets. The proposed model not only improves detection accuracy but also substantially enhances robustness in complex and dynamic traffic scenarios, providing a more reliable solution for autonomous driving in intelligent transportation systems.
No comments yet — start the discussion below.