Civil: causal and intuitive visual imitation learning

Abstract Today’s robots attempt to learn new tasks by imitating human examples. These robots watch the human complete the task, and then try to match the actions taken by the human expert. However, this standard approach to visual imitation learning is fundamentally limited: the robot observes what the human does, but not why the human chooses those behaviors. Without understanding which features of the system or environment factor into the human’s decisions, robot learners often misinterpret the human’s examples (e.g., the robot incorrectly thinks the human picked up a coffee cup because of the color of clutter in the background). In practice, this results in causal confusion, inefficient learning, and robot policies that fail when the environment changes. We therefore propose a shift in perspective: instead of asking human teachers just to show what actions the robot should take, we also enable humans to intuitively indicate why they made those decisions (i.e., what features are critical for the desired task). Under our paradigm human teachers attach markers to task-relevant objects and use natural language prompts to describe their state representation. Our proposed algorithm, CIVIL, leverages this augmented demonstration data to filter the robot’s visual observations and extract a feature representation that aligns with the human teacher. CIVIL then applies these causal features to train a transformer-based policy that — when tested on the robot — is able to emulate human behaviors without being confused by visual distractors or irrelevant items. We show that, by harnessing and combining existing language-grounding, visual-masking, and policy-learning techniques, CIVIL realizes this richer teaching paradigm and incorporates human-provided task relevance directly into imitation learning. Our simulations and real-world experiments demonstrate that robots trained with CIVIL learn both what actions to take and why to take those actions, resulting in better performance than state-of-the-art baselines. From the human’s perspective, our user study reveals that this new training paradigm actually reduces the total time required for the robot to learn the task, and also improves the robot’s performance in previously unseen scenarios. See videos at our project website: https://civil2025.github.io .

Authors

Institutions

Publication Details

Journal
Autonomous Robots
Published
2026-09-18
DOI
https://doi.org/10.1007/s10514-026-10266-3
Primary Topic
Multimodal Machine Learning Applications
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Civil: causal and intuitive visual imitation learning

Yinlong Dai, Dylan P. Losey, Heramb Nemlekar, Shahabedin Sagheb et al.
Autonomous Robots
Multimodal Machine Learning Applications
article

Civil: causal and intuitive visual imitation learning

Yinlong Dai, Dylan P. Losey, Heramb Nemlekar, Shahabedin Sagheb, Robert Ramirez Sanchez, Cara M. Nunez, Ryan Jeronimus
article en

Abstract

Abstract Today’s robots attempt to learn new tasks by imitating human examples. These robots watch the human complete the task, and then try to match the actions taken by the human expert. However, this standard approach to visual imitation learning is fundamentally limited: the robot observes what the human does, but not why the human chooses those behaviors. Without understanding which features of the system or environment factor into the human’s decisions, robot learners often misinterpret the human’s examples (e.g., the robot incorrectly thinks the human picked up a coffee cup because of the color of clutter in the background). In practice, this results in causal confusion, inefficient learning, and robot policies that fail when the environment changes. We therefore propose a shift in perspective: instead of asking human teachers just to show what actions the robot should take, we also enable humans to intuitively indicate why they made those decisions (i.e., what features are critical for the desired task). Under our paradigm human teachers attach markers to task-relevant objects and use natural language prompts to describe their state representation. Our proposed algorithm, CIVIL, leverages this augmented demonstration data to filter the robot’s visual observations and extract a feature representation that aligns with the human teacher. CIVIL then applies these causal features to train a transformer-based policy that — when tested on the robot — is able to emulate human behaviors without being confused by visual distractors or irrelevant items. We show that, by harnessing and combining existing language-grounding, visual-masking, and policy-learning techniques, CIVIL realizes this richer teaching paradigm and incorporates human-provided task relevance directly into imitation learning. Our simulations and real-world experiments demonstrate that robots trained with CIVIL learn both what actions to take and why to take those actions, resulting in better performance than state-of-the-art baselines. From the human’s perspective, our user study reveals that this new training paradigm actually reduces the total time required for the robot to learn the task, and also improves the robot’s performance in previously unseen scenarios. See videos at our project website: https://civil2025.github.io .

Autonomous RobotsVol. 50(4)
California State University, Northridge (US), Cornell University (US), Virginia Tech (US)
National Science Foundation
Openalex Percentile: Top 98%
Multimodal Machine Learning Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.