Vision-encoder-integrated lifting tracker for cost-effective crane operations in modular construction

Cost-effective tracking of modules during lifting remains challenging in modular construction (MC), where existing solutions rely on densely deployed, costly sensors such as Light Detection and Ranging (LiDAR). This paper presents a vision-encoder-integrated lifting tracker (VE-LIFT) that reuses two low-cost on-site assets: a monocular camera and the crane's factory-installed encoders. The camera estimates a cylindrical bounding volume (CBV) enveloping the module through a three-stage pipeline: a You Only Look Once (YOLO)v8-Pose model detects the module, lifting frame, and keypoints; Segment Anything Model (SAM) 3 refines keypoints via box and text prompts; and a hierarchical Perspective-n-Point solver derives the CBV dimensions. Encoders update the CBV position via Modbus-over-Ethernet. On a real-life MC project, VE-LIFT achieved 98.65%/98.93% [email protected]:0.95 for box/pose, over 22% RMSE reduction by SAM 3, and sub-0.5 m tracking during module lifting, suggesting LiDAR-comparable accuracy at lower cost. VE-LIFT delivers AI-based high-accuracy and cost-effective lifting tracking without additional dedicated sensors.

Authors

Institutions

Publication Details

Journal
Automation in Construction
Published
2026-09-18
DOI
https://doi.org/10.1016/j.autcon.2026.107262
Primary Topic
BIM and Construction Integration
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Vision-encoder-integrated lifting tracker for cost-effective crane operations in modular construction

Jiayi Xu, Aimin Zhu, Wei Pan, Qiqi Zhang et al.
Automation in Construction
BIM and Construction Integration
article

Vision-encoder-integrated lifting tracker for cost-effective crane operations in modular construction

Jiayi Xu, Aimin Zhu, Wei Pan, Qiqi Zhang, Zhiqian Zhang, Kai Wang
article en

Abstract

Cost-effective tracking of modules during lifting remains challenging in modular construction (MC), where existing solutions rely on densely deployed, costly sensors such as Light Detection and Ranging (LiDAR). This paper presents a vision-encoder-integrated lifting tracker (VE-LIFT) that reuses two low-cost on-site assets: a monocular camera and the crane's factory-installed encoders. The camera estimates a cylindrical bounding volume (CBV) enveloping the module through a three-stage pipeline: a You Only Look Once (YOLO)v8-Pose model detects the module, lifting frame, and keypoints; Segment Anything Model (SAM) 3 refines keypoints via box and text prompts; and a hierarchical Perspective-n-Point solver derives the CBV dimensions. Encoders update the CBV position via Modbus-over-Ethernet. On a real-life MC project, VE-LIFT achieved 98.65%/98.93% [email protected]:0.95 for box/pose, over 22% RMSE reduction by SAM 3, and sub-0.5 m tracking during module lifting, suggesting LiDAR-comparable accuracy at lower cost. VE-LIFT delivers AI-based high-accuracy and cost-effective lifting tracking without additional dedicated sensors.

Automation in ConstructionVol. 192
University of Hong Kong (HK)
Research Grants Council, University Grants Committee, Food and Health Bureau
Industry, innovation and infrastructure
Openalex Percentile: Top 15%
BIM and Construction Integration
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Vision-encoder-integrated lifting tracker for cost-effective crane operations in modular construction — Jiayi Xu, Aimin Zhu, et al. · Automation in Construction (2026) | TGRS Research Map | TGRS