AlignDual : Real-Time Multimodal 3D Detection on Resource-Constraint Edge Device via Inter-Stream Cooperation

Multimodal 3D Object Detection (M3DOD) is critical for applications like city surveillance, industrial defect detection, and autonomous driving, yet state-of-the-art algorithms are too resource-intensive for edge devices. While cloud offloading appears to be a solution, we identify a fundamental data misalignment problem inherent to this approach that degrades detection performance. This failure manifests as two intertwined issues: (1) Semantic Misalignment , where uncoordinated, modality-agnostic compression can discard the cross-modal spatial correlations essential for fusion, and (2) Temporal Misalignment , where heterogeneous pipeline delays lead to substantial synchronization bottlenecks and violate real-time constraints. This paper introduces AlignDual , a novel framework that tackles these challenges by establishing a new paradigm: Cross-Modal Co-Design. Instead of treating sensor streams as independent flows, AlignDual establishes two key cooperative mechanisms. First, a semantically-coordinated compression scheme leverages edge-efficient 2D object semantics to guide point cloud sampling at the source, preserving fusion-critical correlations before transmission. Second, a proactive, prediction-based synchronization framework abandons reactive waiting, instead using motion prediction to compensate for latency jitter and reduce synchronization overhead. These mechanisms are orchestrated by a closed-loop optimizer that dynamically adapts to runtime conditions. We implemented and evaluated AlignDual on a real-world testbed. Results show that our system outperforms state-of-the-art cloud-based approaches, improving detection accuracy (mAP) by up to 18.2% while simultaneously increasing the real-time latency compliance rate (CR) by 22.3%.

Authors

Institutions

Publication Details

Journal
Proceedings of the ACM on Interactive Mobile Wearable and Ubiquitous Technologies
Published
2026-09-30
DOI
https://doi.org/10.1145/3831654
Primary Topic
Advanced Neural Network Applications
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

AlignDual : Real-Time Multimodal 3D Detection on Resource-Constraint Edge Device via Inter-Stream Cooperation

Bin Qian, Shibo He, Liping Qian, Zhenyu Wen et al.
Proceedings of the ACM on Interactive Mobile Wearable and Ubiquitous Technologies
Advanced Neural Network Applications
article

AlignDual : Real-Time Multimodal 3D Detection on Resource-Constraint Edge Device via Inter-Stream Cooperation

Bin Qian, Shibo He, Liping Qian, Zhenyu Wen, Cong Wang, Zhen Hong, Xiaoli Zhang, Tao Wang, Zihua Yang, Yong Zhu
article en

Abstract

Multimodal 3D Object Detection (M3DOD) is critical for applications like city surveillance, industrial defect detection, and autonomous driving, yet state-of-the-art algorithms are too resource-intensive for edge devices. While cloud offloading appears to be a solution, we identify a fundamental data misalignment problem inherent to this approach that degrades detection performance. This failure manifests as two intertwined issues: (1) Semantic Misalignment , where uncoordinated, modality-agnostic compression can discard the cross-modal spatial correlations essential for fusion, and (2) Temporal Misalignment , where heterogeneous pipeline delays lead to substantial synchronization bottlenecks and violate real-time constraints. This paper introduces AlignDual , a novel framework that tackles these challenges by establishing a new paradigm: Cross-Modal Co-Design. Instead of treating sensor streams as independent flows, AlignDual establishes two key cooperative mechanisms. First, a semantically-coordinated compression scheme leverages edge-efficient 2D object semantics to guide point cloud sampling at the source, preserving fusion-critical correlations before transmission. Second, a proactive, prediction-based synchronization framework abandons reactive waiting, instead using motion prediction to compensate for latency jitter and reduce synchronization overhead. These mechanisms are orchestrated by a closed-loop optimizer that dynamically adapts to runtime conditions. We implemented and evaluated AlignDual on a real-world testbed. Results show that our system outperforms state-of-the-art cloud-based approaches, improving detection accuracy (mAP) by up to 18.2% while simultaneously increasing the real-time latency compliance rate (CR) by 22.3%.

Proceedings of the ACM on Interactive Mobile Wearable and Ubiquitous TechnologiesVol. 10(3)
Zhejiang University of Technology (CN), Zhejiang University (CN), University of Science and Technology Beijing (CN)
Openalex Percentile: Top 14%
Advanced Neural Network Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.