Jiawei Sun · Applied and Computational Engineering 2026 · 2026
DOI: 10.54254/2755-2721/2026.36700
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
With the large-scale implementation of scenarios such as logistics warehousing and park inspections, the demand for autonomous environmental perception by mobile robots continues to grow. Low-cost pure vision semantic perception has become a core technology for ensuring autonomous safe navigation of these robots. Traditional manual feature algorithms lack stability in complex road conditions and varying lighting conditions, while deep learning relies on autonomous hierarchical feature learning to break through the original performance limit. This article focuses on real-time semantic perception of embedded mobile cars, systematically sorting out three mainstream deep learning solutions: CNN, lightweight network, and visual Transformer, and comparing the adaptation differences of each architecture in segmentation accuracy, inference speed, and complex working conditions. Summarize the existing technological bottlenecks from three aspects: imbalanced real-time accuracy, weak robustness in complex environments, and difficulty in recognizing distant and occluded targets. Corresponding optimization ideas are proposed, including dynamic lightweight inference, cross-domain adaptive enhancement, convolutional attention hybrid architecture, and end-to-end collaborative deployment. The article comprehensively outlines the development trajectory of visual semantic perception technology for mobile vehicles, providing a systematic theoretical reference for the design and embedded deployment optimization of visual algorithms for mobile robots in warehousing and inspection tasks.
No comments yet — start the discussion below.