AlignDual : Real-Time Multimodal 3D Detection on Resource-Constraint Edge Device via Inter-Stream Cooperation
Multimodal 3D Object Detection (M3DOD) is critical for applications like city surveillance, industrial defect detection, and autonomous driving, yet state-of-the-art algorithms are too resource-intensive for edge devices. While cloud offloading appears to be a solution, we identify a fundamental data misalignment problem inherent to this approach that degrades detection performance. This failure manifests as two intertwined issues: (1) Semantic Misalignment , where uncoordinated, modality-agnostic compression can discard the cross-modal spatial correlations essential for fusion, and (2) Temporal Misalignment , where heterogeneous pipeline delays lead to substantial synchronization bottlenecks and violate real-time constraints. This paper introduces AlignDual , a novel framework that tackles these challenges by establishing a new paradigm: Cross-Modal Co-Design. Instead of treating sensor streams as independent flows, AlignDual establishes two key cooperative mechanisms. First, a semantically-coordinated compression scheme leverages edge-efficient 2D object semantics to guide point cloud sampling at the source, preserving fusion-critical correlations before transmission. Second, a proactive, prediction-based synchronization framework abandons reactive waiting, instead using motion prediction to compensate for latency jitter and reduce synchronization overhead. These mechanisms are orchestrated by a closed-loop optimizer that dynamically adapts to runtime conditions. We implemented and evaluated AlignDual on a real-world testbed. Results show that our system outperforms state-of-the-art cloud-based approaches, improving detection accuracy (mAP) by up to 18.2% while simultaneously increasing the real-time latency compliance rate (CR) by 22.3%.
Authors
- Bin Qian (ORCID: https://orcid.org/0000-0001-7058-0360)
- Shibo He (ORCID: https://orcid.org/0000-0002-1505-6766)
- Liping Qian (ORCID: https://orcid.org/0000-0001-6210-2617)
- Zhenyu Wen (ORCID: https://orcid.org/0000-0002-2914-912X)
- Cong Wang (ORCID: https://orcid.org/0009-0008-4818-3797)
- Zhen Hong (ORCID: https://orcid.org/0000-0001-9956-3732)
- Xiaoli Zhang (ORCID: https://orcid.org/0009-0005-8317-5539)
- Tao Wang (ORCID: https://orcid.org/0009-0002-1206-2201)
- Zihua Yang (ORCID: https://orcid.org/0009-0007-9894-4462)
- Yong Zhu (ORCID: https://orcid.org/0009-0005-6504-8579)
Institutions
- Zhejiang University of Technology (CN)
- Zhejiang University (CN)
- University of Science and Technology Beijing (CN)
Publication Details
- Journal
- Proceedings of the ACM on Interactive Mobile Wearable and Ubiquitous Technologies
- Published
- 2026-09-30
- DOI
- https://doi.org/10.1145/3831654
- Primary Topic
- Advanced Neural Network Applications
- Type
- article
- Field-Weighted Citation Impact
- 0.00