Predictive Visual Representations for Generalizable Vision-Based Navigation

This project investigates whether action-conditioned future latent prediction can improve visual representations learned by a vision-based navigation policy and their transfer to downstream reinforcement learning. The study uses the MetaDrive simulator to train and evaluate vision-based driving agents on procedurally generated road layouts, comparing behaviour cloning with and without auxiliary representation-learning objectives. The project includes controlled ablations of temporal prediction, action conditioning, and same-timestep cross-view invariance, together with downstream PPO transfer and representation analysis.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-28
DOI
https://doi.org/10.5281/zenodo.23009956
Primary Topic
Autonomous Vehicle Technology and Safety
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Predictive Visual Representations for Generalizable Vision-Based Navigation

Shubham Waghmare
Zenodo (CERN European Organization for Nuclear Research)
Autonomous Vehicle Technology and Safety
preprint

Predictive Visual Representations for Generalizable Vision-Based Navigation

Shubham Waghmare
preprint en

Abstract

This project investigates whether action-conditioned future latent prediction can improve visual representations learned by a vision-based navigation policy and their transfer to downstream reinforcement learning. The study uses the MetaDrive simulator to train and evaluate vision-based driving agents on procedurally generated road layouts, comparing behaviour cloning with and without auxiliary representation-learning objectives. The project includes controlled ablations of temporal prediction, action conditioning, and same-timestep cross-view invariance, together with downstream PPO transfer and representation analysis.

Zenodo (CERN European Organization for Nuclear Research)
Autonomous Vehicle Technology and Safety
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Predictive Visual Representations for Generalizable Vision-Based Navigation — Shubham Waghmare · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS