Xiang Ji, Shuaifei Hu, Haofei Wang, Wanming Hao, Xiangnan Li · Electronics 2026 · 2026
DOI: 10.3390/electronics15194431
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Efficient deployment of neural-network detectors on field-programmable gate array (FPGA) accelerators depends not only on model complexity but also on compiler-visible graph structure. On deep learning processing unit (DPU) platforms, unsupported operators can fragment execution between accelerator and host execution domains. We present C2fDeploy, a training-free rewrite for the YOLOv8 C2f block that moves channel splitting from the activation graph to an offline partition of the trained projection and batch-normalization parameters. The transformed block preserves the 32-bit floating-point (FP32) function without retraining or additional parameters. On the GRAZPEDWRI-DX fracture-detection task, the original and rewritten graphs produced identical FP32 test metrics and comparable 8-bit integer (INT8) accuracy. Direct tensor-level FP32 comparison at the outputs of all eight rewritten C2f blocks yielded an aggregate mean absolute error of 9.73×10−8 and a relative L2 error of 1.97×10−7, providing numerical verification beyond detection-level metrics. Compilation for the Kria KV260 consolidated nine DPU subgraphs into one and removed the C2f-related host-side slicing operations. In a same-checkpoint whole-XModel benchmark, C2fDeploy improved whole-XModel graph execution throughput by 95.2× and reduced energy per execution on the 5 V system-on-module (SOM) rail by 98.2%. Direct runtime profiling further showed that DPU compute-unit busy-time utilization increased from 0.36% to 90.75%, while system-wide CPU utilization decreased by 89.5%. Aggregate APM-observed external-memory bandwidth increased from 40.20 to 3250.56 MB/s as accelerator execution became more continuous, whereas normalized APM-observed traffic decreased by 14.8% per graph execution. These results show that compiler-aware graph rewriting can remove deployment bottlenecks without changing the trained detector.
No comments yet — start the discussion below.