
Yi Wang, Meng Zhang, Yong Zhang, Xia Sheng, Xiaohua Cheng · Engineering Research Express 2026 · 2026
DOI: 10.1088/2631-8695/aeafba
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Visual environmental perception is the safety cornerstone of autonomous driving, yet robust vehicle detection remains challenging under intense glare, occlusion, and adverse weather. Although YOLOv11 provides an efficient baseline, it suffers from coupled frequency-spatial extraction and inadequate long-range dependency modeling, leading to misdetections of occluded or distant targets. To circumvent these bottlenecks, we propose a collaborative enhanced YOLOv11 framework integrating three innovative modules: a Frequency-Spatial Decoupled Feature Pyramid Network (FSD-FPN), a Visual Recursive State-Space Module (VRSSM), and a Multi-scale Attention Convolution-Transformer Block (MACTB). Specifically, FSD-FPN segregates high-frequency structural details from low-frequency semantic components to attenuate spectral noise induced by extreme illumination . The Mamba-based VRSSM establishes global contextual dependencies with linear complexity to reconstruct feature integrity for occluded vehicles , while MACTB synergizes local convolutional biases with global Transformer refinement to enhance multi-scale discriminative power . Empirical evaluations on the BDD100K and a self-curated Chengdu urban dataset demonstrate that our framework achieves a superior mean Average Precision (mAP)@0.5:0.95 of 70.9\% and an mAP@0.5 of 93.0\%, consistently outperforming YOLOv7, YOLOv10, vanilla YOLOv11, and other state-of-the-art detectors. With a throughput of 173 FPS and a minimal memory footprint of 5.2 MB, the model ensures exceptional efficiency on resource-constrained edge devices , providing a robust and industrially viable solution for autonomous vehicle perception in adverse environments.
No comments yet — start the discussion below.