Na Zhang, Edmundo Guerra, Antoni Grau · Electronics 2026 · 2026
DOI: 10.3390/electronics15184200
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Multi-modal 3D object detection is critical for autonomous driving perception. While Bird’s Eye View (BEV) fusion methods effectively integrate LiDAR and camera features, they primarily focus on single-frame fusion and neglect temporal context. We propose CamT-BEV, a camera-temporal-enhanced BEV fusion framework for improved multi-modal 3D object detection. Our key insight is that temporal modeling is particularly critical for the camera branch to resolve monocular depth ambiguity and object occlusion, while single-frame LiDAR representation already provides accurate instantaneous geometry. We thus propose a camera-centric temporal enhancement module via ego-motion warping and ConvLSTM temporal encoding. Extensive experiments on the nuScenes dataset demonstrate that CamT-BEV achieves competitive perception performance, attaining 0.6971 NDS and 0.6683 mAP, with notable relative AP gains on challenging categories such as bicycles (+27.3%) and motorcycles (+7.66%) evaluated under category-level mAP (averaged across 0.5 m to 4.0 m distance thresholds). Furthermore, evaluations under fog and miss-beam conditions in nuScenes-C confirm its improved robustness against specific visual and sensor degradations. Crucially, these gains are achieved with low additional computational and memory overhead, demonstrating that targeted camera-temporal fusion is a practical solution for 3D perception.
No comments yet — start the discussion below.