A vision-language conditioned agent for autonomous excavator operations in physics-based simulation
This study presents a vision-language conditioned agent for autonomous excavator operations in physics-based simulation. A semantic action-based instruction scheme is introduced, in which natural language instructions describe excavator action units with target specifications. The agent integrates a vision-language model (VLM) with a reinforcement learning policy to generate continuous actuator-level commands from semantic instructions, visual observations, and kinematic state. A hierarchical reward combining a terminal goal term with step-wise motion shaping and stability terms aligns learning with semantic correctness, embodiment-level progression, and mechanical admissibility. Experiments in physics-based simulation show that the agent achieves 92.4 % average success across seven semantic actions, with ablations confirming the contribution of the hierarchical reward and the joint multimodal grounding. The agent reaches dig-dump success rates of 83.0 % and 81.0 % under fixed and randomized targets, outperforming recent methods. In multi-cycle evaluation, the agent completes area excavation and trenching tasks without additional training by sequentially composing the learned semantic actions, with productivities comparable to earthmoving practice. All these tasks are executed through instruction input alone, enabling simulation-driven analyses toward improved productivity and safety in construction.
Authors
- Seunghoon Jung (ORCID: https://orcid.org/0000-0001-8913-6390)
- Taehoon Hong (ORCID: https://orcid.org/0000-0001-5136-8276)
Institutions
- Yonsei University (KR)
Publication Details
- Journal
- Advanced Engineering Informatics
- Published
- 2026-09-18
- DOI
- https://doi.org/10.1016/j.aei.2026.105287
- Primary Topic
- Hydraulic and Pneumatic Systems
- Type
- article
- Field-Weighted Citation Impact
- 0.00
Funders
- National Research Foundation of Korea