XIE Jingyuan, ZHA Kaiwen, LIU Pengju, TIAN Chunwei · DOAJ (DOAJ: Directory of Open Access Journals) 2026 · 2026
DOI: 10.19678/j.issn.1000-3428.0260413
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
This paper studies operator-level reconstruction for the structural mismatch between YOLOv11 and Ascend Neural Processing Unit (NPU). The Spatial Pyramid Pooling-Fast (SPPF), C3K2, and C2PSA modules are optimized without changing network semantics or model scale. Three Ascend C operators are designed: the SPPF operator uses on-chip data loop and halo cache to reduce redundant global-memory traffic in multi-stage pooling; the C3K2 operator integrates multi-core task assignment and multi-queue asynchronous pipelining to reduce fine-grained kernel launch overhead; and the C2PSA operator reconstructs attention communication through a parallel reduction-broadcast primitive. On an Ascend 910B NPU, the complete reconstruction reduces the training time per epoch by 23.2% and improves the training throughput by 27.6% on the COCO dataset. The results show that matching Ascend on-chip memory, asynchronous queues, and multi-core synchronization mechanisms improves the training execution efficiency of key YOLOv11 modules and keeps inference performance stable. It can provide verifiable operator mapping schemes for the deployment of complex object detection networks on the Ascend platform.
No comments yet — start the discussion below.