Temporal Memory for Robust Navigation under Distribution Shift: A Controlled Benchmark of Recurrent and Memoryless Policies

This release contains the complete code, trained models, evaluation scripts, and paper source for a controlled benchmark studying the role of temporal memory in robust mobile robot navigation under distribution shift. MOTIVATIONDeep reinforcement learning has shown promise for learning navigation policies directly from sensor input, but these policies often fail when deployed outside their training distribution. Domain randomization is a standard technique for bridging this gap, yet the specific contribution of temporal memory to robustness has not been systematically isolated under controlled conditions. This work addresses that gap. APPROACHWe compare a memoryless MLP policy trained with Proximal Policy Optimization (PPO) against a recurrent LSTM policy trained with RecurrentPPO on a goal-occlusion navigation task. In this task, an obstacle periodically blocks the line of sight between the robot and the goal, requiring the policy to remember the last visible goal direction. Both policies are trained with mild domain randomization and evaluated under three conditions of increasing difficulty: in-distribution (no domain randomization), the training distribution, and an unseen severe distribution shift with additional obstacles, faster dynamics, heavier sensor noise, and longer action latency. KEY RESULTUnder severe distribution shift, the recurrent LSTM policy degrades by only 8 percentage points from its training-distribution success rate, while the memoryless MLP collapses by 30 points. The LSTM also learns faster and reaches a higher asymptotic reward. However, both policies fail to solve the severe condition, with success rates below 5%, indicating that temporal memory is a necessary but not sufficient ingredient for robust navigation in this class of tasks. CONTENTS- Complete training and evaluation pipeline (PPO and RecurrentPPO)- Lightweight 2D navigation environment with goal occlusion, implemented in pure NumPy for CPU-only execution- Trained model checkpoints and all evaluation results in CSV format- Paper source (LaTeX) and compiled PDF- Reproducibility scripts for all figures REQUIREMENTSPython 3.11, PyTorch 2.x (CPU), Stable-Baselines3, sb3-contrib.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-09
DOI
https://doi.org/10.5281/zenodo.23256937
Primary Topic
Reinforcement Learning in Robotics
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Temporal Memory for Robust Navigation under Distribution Shift: A Controlled Benchmark of Recurrent and Memoryless Policies

Sultan Ali Khan
Zenodo (CERN European Organization for Nuclear Research)
Reinforcement Learning in Robotics
article

Temporal Memory for Robust Navigation under Distribution Shift: A Controlled Benchmark of Recurrent and Memoryless Policies

Sultan Ali Khan
article en

Abstract

This release contains the complete code, trained models, evaluation scripts, and paper source for a controlled benchmark studying the role of temporal memory in robust mobile robot navigation under distribution shift. MOTIVATIONDeep reinforcement learning has shown promise for learning navigation policies directly from sensor input, but these policies often fail when deployed outside their training distribution. Domain randomization is a standard technique for bridging this gap, yet the specific contribution of temporal memory to robustness has not been systematically isolated under controlled conditions. This work addresses that gap. APPROACHWe compare a memoryless MLP policy trained with Proximal Policy Optimization (PPO) against a recurrent LSTM policy trained with RecurrentPPO on a goal-occlusion navigation task. In this task, an obstacle periodically blocks the line of sight between the robot and the goal, requiring the policy to remember the last visible goal direction. Both policies are trained with mild domain randomization and evaluated under three conditions of increasing difficulty: in-distribution (no domain randomization), the training distribution, and an unseen severe distribution shift with additional obstacles, faster dynamics, heavier sensor noise, and longer action latency. KEY RESULTUnder severe distribution shift, the recurrent LSTM policy degrades by only 8 percentage points from its training-distribution success rate, while the memoryless MLP collapses by 30 points. The LSTM also learns faster and reaches a higher asymptotic reward. However, both policies fail to solve the severe condition, with success rates below 5%, indicating that temporal memory is a necessary but not sufficient ingredient for robust navigation in this class of tasks. CONTENTS- Complete training and evaluation pipeline (PPO and RecurrentPPO)- Lightweight 2D navigation environment with goal occlusion, implemented in pure NumPy for CPU-only execution- Trained model checkpoints and all evaluation results in CSV format- Paper source (LaTeX) and compiled PDF- Reproducibility scripts for all figures REQUIREMENTSPython 3.11, PyTorch 2.x (CPU), Stable-Baselines3, sb3-contrib.

Zenodo (CERN European Organization for Nuclear Research)
Openalex Percentile: Top 12%
Reinforcement Learning in Robotics
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Temporal Memory for Robust Navigation under Distribution Shift: A Controlled Benchmark of Recurrent and Memoryless Policies — Sultan Ali Khan · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS