Affordance-based robot manipulation with flow matching
We present a framework for assistive robot manipulation that addresses two fundamental challenges: efficient adaptation of large-scale models for scene affordance understanding and effective learning of robot actions by grounding the visual affordance. To tackle the first challenge, we adopt a parameter-efficient prompt tuning method, prepending learnable text prompts to a frozen vision model to predict affordances, while considering spatial and semantic relationships in multi-task scenarios. For the second challenge, we propose a flow matching method, representing a robot visuomotor policy as a conditional process of flowing random waypoints to desired robot actions. We introduce a real-world dataset with 10 tasks to evaluate our approach. Experiments show our prompt tuning method achieves competitive or superior performance to other finetuning protocols across data scales, while satisfying parameter efficiency. Furthermore, flow matching yields more stable training and faster inference compared to diffusion policy; specifically, 1-step flow matching achieves comparable accuracy to 16-step DDIM while reducing inference time by roughly 85%. Our framework seamlessly unifies high-level parameter-efficient affordance representation learning with low-level flow matching policies, illustrating how explicit affordances serve as effective spatial grounding for flow-based policies. https://github.com/HRI-EU/flow_matching .
Authors
- Michael Gienger (ORCID: https://orcid.org/0000-0001-8036-2519)
- Fan Zhang (ORCID: https://orcid.org/0009-0002-7980-1232)
Institutions
- Honda (Japan) (JP)
Publication Details
- Journal
- Frontiers in Robotics and AI
- Published
- 2026-09-22
- DOI
- https://doi.org/10.3389/frobt.2026.1906861
- Citations
- 1
- Primary Topic
- Reinforcement Learning in Robotics
- Type
- article
- Field-Weighted Citation Impact
- 0.00