Predictive Visual Representations for Generalizable Vision-Based Navigation
This project investigates whether action-conditioned future latent prediction can improve visual representations learned by a vision-based navigation policy and their transfer to downstream reinforcement learning. The study uses the MetaDrive simulator to train and evaluate vision-based driving agents on procedurally generated road layouts, comparing behaviour cloning with and without auxiliary representation-learning objectives. The project includes controlled ablations of temporal prediction, action conditioning, and same-timestep cross-view invariance, together with downstream PPO transfer and representation analysis.
Authors
- Shubham Waghmare
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-28
- DOI
- https://doi.org/10.5281/zenodo.23009956
- Primary Topic
- Autonomous Vehicle Technology and Safety
- Type
- preprint