A vision-language conditioned agent for autonomous excavator operations in physics-based simulation

This study presents a vision-language conditioned agent for autonomous excavator operations in physics-based simulation. A semantic action-based instruction scheme is introduced, in which natural language instructions describe excavator action units with target specifications. The agent integrates a vision-language model (VLM) with a reinforcement learning policy to generate continuous actuator-level commands from semantic instructions, visual observations, and kinematic state. A hierarchical reward combining a terminal goal term with step-wise motion shaping and stability terms aligns learning with semantic correctness, embodiment-level progression, and mechanical admissibility. Experiments in physics-based simulation show that the agent achieves 92.4 % average success across seven semantic actions, with ablations confirming the contribution of the hierarchical reward and the joint multimodal grounding. The agent reaches dig-dump success rates of 83.0 % and 81.0 % under fixed and randomized targets, outperforming recent methods. In multi-cycle evaluation, the agent completes area excavation and trenching tasks without additional training by sequentially composing the learned semantic actions, with productivities comparable to earthmoving practice. All these tasks are executed through instruction input alone, enabling simulation-driven analyses toward improved productivity and safety in construction.

Authors

Institutions

Publication Details

Journal
Advanced Engineering Informatics
Published
2026-09-18
DOI
https://doi.org/10.1016/j.aei.2026.105287
Primary Topic
Hydraulic and Pneumatic Systems
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

A vision-language conditioned agent for autonomous excavator operations in physics-based simulation

Seunghoon Jung, Taehoon Hong
Advanced Engineering Informatics
Hydraulic and Pneumatic Systems
article

A vision-language conditioned agent for autonomous excavator operations in physics-based simulation

Seunghoon Jung, Taehoon Hong
article en

Abstract

This study presents a vision-language conditioned agent for autonomous excavator operations in physics-based simulation. A semantic action-based instruction scheme is introduced, in which natural language instructions describe excavator action units with target specifications. The agent integrates a vision-language model (VLM) with a reinforcement learning policy to generate continuous actuator-level commands from semantic instructions, visual observations, and kinematic state. A hierarchical reward combining a terminal goal term with step-wise motion shaping and stability terms aligns learning with semantic correctness, embodiment-level progression, and mechanical admissibility. Experiments in physics-based simulation show that the agent achieves 92.4 % average success across seven semantic actions, with ablations confirming the contribution of the hierarchical reward and the joint multimodal grounding. The agent reaches dig-dump success rates of 83.0 % and 81.0 % under fixed and randomized targets, outperforming recent methods. In multi-cycle evaluation, the agent completes area excavation and trenching tasks without additional training by sequentially composing the learned semantic actions, with productivities comparable to earthmoving practice. All these tasks are executed through instruction input alone, enabling simulation-driven analyses toward improved productivity and safety in construction.

Advanced Engineering InformaticsVol. 77
Yonsei University (KR)
National Research Foundation of Korea
Decent work and economic growth
Openalex Percentile: Top 20%
Hydraulic and Pneumatic Systems
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

A vision-language conditioned agent for autonomous excavator operations in physics-based simulation — Seunghoon Jung, Taehoon Hong · Advanced Engineering Informatics (2026) | TGRS Research Map | TGRS