Temporal Memory for Robust Navigation under Distribution Shift: A Controlled Benchmark of Recurrent and Memoryless Policies
This release contains the complete code, trained models, evaluation scripts, and paper source for a controlled benchmark studying the role of temporal memory in robust mobile robot navigation under distribution shift. MOTIVATIONDeep reinforcement learning has shown promise for learning navigation policies directly from sensor input, but these policies often fail when deployed outside their training distribution. Domain randomization is a standard technique for bridging this gap, yet the specific contribution of temporal memory to robustness has not been systematically isolated under controlled conditions. This work addresses that gap. APPROACHWe compare a memoryless MLP policy trained with Proximal Policy Optimization (PPO) against a recurrent LSTM policy trained with RecurrentPPO on a goal-occlusion navigation task. In this task, an obstacle periodically blocks the line of sight between the robot and the goal, requiring the policy to remember the last visible goal direction. Both policies are trained with mild domain randomization and evaluated under three conditions of increasing difficulty: in-distribution (no domain randomization), the training distribution, and an unseen severe distribution shift with additional obstacles, faster dynamics, heavier sensor noise, and longer action latency. KEY RESULTUnder severe distribution shift, the recurrent LSTM policy degrades by only 8 percentage points from its training-distribution success rate, while the memoryless MLP collapses by 30 points. The LSTM also learns faster and reaches a higher asymptotic reward. However, both policies fail to solve the severe condition, with success rates below 5%, indicating that temporal memory is a necessary but not sufficient ingredient for robust navigation in this class of tasks. CONTENTS- Complete training and evaluation pipeline (PPO and RecurrentPPO)- Lightweight 2D navigation environment with goal occlusion, implemented in pure NumPy for CPU-only execution- Trained model checkpoints and all evaluation results in CSV format- Paper source (LaTeX) and compiled PDF- Reproducibility scripts for all figures REQUIREMENTSPython 3.11, PyTorch 2.x (CPU), Stable-Baselines3, sb3-contrib.
Authors
- Sultan Ali Khan
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-09
- DOI
- https://doi.org/10.5281/zenodo.23256937
- Primary Topic
- Reinforcement Learning in Robotics
- Type
- article
- Field-Weighted Citation Impact
- 0.00