Shuqi Wang, Xinyi Zhang · Electronics 2026 · 2026
DOI: 10.3390/electronics15184123
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
This study addresses the challenges of sparse long-range point clouds, complex background interference, and inconsistent localization quality in 3D object detection for intelligent driving in open-pit mines. A camera–LiDAR multimodal detection method based on Voxel R-CNN is proposed. We introduce a Multimodal Focal Sparse Voxel Enhancement module that combines Focal Sparse Convolution with shallow visual features to guide voxel importance prediction and selective sparse propagation. An Intersection over Union (IoU)-Aware Quality and Geometry Refinement Head is further designed to improve the localization accuracy and ranking reliability of 3D proposals. Experimental results show that the proposed method achieves BEV mAP@0.40 and 3D mAP@0.40 values of 63.43% and 61.23%, respectively, outperforming the strongest comparison method, MambaFusion, by 5.37 and 7.61 percentage points. In the 60–80 m range, the average translation error is reduced to 0.41 m, while the inference speed reaches 27 FPS. These results demonstrate that the proposed method improves the detection and localization of distant sparse objects in complex open-pit mine environments while maintaining real-time inference capability.
No comments yet — start the discussion below.