chiheb nouri · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22996765
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
This paper presents a decoupled Producer–Consumer architecture for low-latency enterprise video analytics on legacy edge hardware. The proposed architecture separates network video ingestion from AI inference, preventing network buffer buildup and allowing the inference pipeline to continuously process the freshest available frame. The system integrates RF-DETR object detection, ONNX Runtime with CUDA acceleration, RTSP/FFmpeg ingestion, WebRTC delivery through Mediasoup, and real-time people counting and object tracking. The architecture was evaluated on a legacy enterprise server equipped with dual Intel Xeon Gold 5120 CPUs, 126 GB RAM, and an NVIDIA Quadro M4000 GPU. The proposed design achieved a stable 13.0 FPS inference rate with approximately 39.4 ms GPU inference time while addressing practical production issues including FFmpeg congestion, OpenCV memory-management failures, WebRTC state synchronization, reverse-proxy buffering, and restrictive enterprise network environments. The results demonstrate how software architecture and asynchronous pipeline design can significantly extend the useful lifetime of existing edge infrastructure for real-time AI video analytics.
No comments yet — start the discussion below.