Yong Zhu, Zhenyu Wen, Tao Wang, Zihua Yang, Xiaoli Zhang, Zhen Hong, Bin Qian, Shibo He, Liping Qian, Cong Wang · Proceedings of the ACM on Interactive Mobile Wearable and Ubiquitous Technologies 2026 · 2026
DOI: 10.1145/3831654
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Multimodal 3D Object Detection (M3DOD) is critical for applications like city surveillance, industrial defect detection, and autonomous driving, yet state-of-the-art algorithms are too resource-intensive for edge devices. While cloud offloading appears to be a solution, we identify a fundamental data misalignment problem inherent to this approach that degrades detection performance. This failure manifests as two intertwined issues: (1) Semantic Misalignment , where uncoordinated, modality-agnostic compression can discard the cross-modal spatial correlations essential for fusion, and (2) Temporal Misalignment , where heterogeneous pipeline delays lead to substantial synchronization bottlenecks and violate real-time constraints. This paper introduces AlignDual , a novel framework that tackles these challenges by establishing a new paradigm: Cross-Modal Co-Design. Instead of treating sensor streams as independent flows, AlignDual establishes two key cooperative mechanisms. First, a semantically-coordinated compression scheme leverages edge-efficient 2D object semantics to guide point cloud sampling at the source, preserving fusion-critical correlations before transmission. Second, a proactive, prediction-based synchronization framework abandons reactive waiting, instead using motion prediction to compensate for latency jitter and reduce synchronization overhead. These mechanisms are orchestrated by a closed-loop optimizer that dynamically adapts to runtime conditions. We implemented and evaluated AlignDual on a real-world testbed. Results show that our system outperforms state-of-the-art cloud-based approaches, improving detection accuracy (mAP) by up to 18.2% while simultaneously increasing the real-time latency compliance rate (CR) by 22.3%.
No comments yet — start the discussion below.