Vision-encoder-integrated lifting tracker for cost-effective crane operations in modular construction
Cost-effective tracking of modules during lifting remains challenging in modular construction (MC), where existing solutions rely on densely deployed, costly sensors such as Light Detection and Ranging (LiDAR). This paper presents a vision-encoder-integrated lifting tracker (VE-LIFT) that reuses two low-cost on-site assets: a monocular camera and the crane's factory-installed encoders. The camera estimates a cylindrical bounding volume (CBV) enveloping the module through a three-stage pipeline: a You Only Look Once (YOLO)v8-Pose model detects the module, lifting frame, and keypoints; Segment Anything Model (SAM) 3 refines keypoints via box and text prompts; and a hierarchical Perspective-n-Point solver derives the CBV dimensions. Encoders update the CBV position via Modbus-over-Ethernet. On a real-life MC project, VE-LIFT achieved 98.65%/98.93% [email protected]:0.95 for box/pose, over 22% RMSE reduction by SAM 3, and sub-0.5 m tracking during module lifting, suggesting LiDAR-comparable accuracy at lower cost. VE-LIFT delivers AI-based high-accuracy and cost-effective lifting tracking without additional dedicated sensors.
Authors
- Jiayi Xu
- Aimin Zhu (ORCID: https://orcid.org/0000-0003-4580-1090)
- Wei Pan
- Qiqi Zhang
- Zhiqian Zhang
- Kai Wang
Institutions
- University of Hong Kong (HK)
Publication Details
- Journal
- Automation in Construction
- Published
- 2026-09-18
- DOI
- https://doi.org/10.1016/j.autcon.2026.107262
- Primary Topic
- BIM and Construction Integration
- Type
- article
- Field-Weighted Citation Impact
- 0.00
Funders
- Research Grants Council, University Grants Committee
- Food and Health Bureau