PV-STAM: Velocity-Aware Attention for Mapless Deep Reinforcement Learning Navigation in Dynamic Environments
Mapless deep reinforcement learning (DRL) navigation in dynamic indoor environments is difficult under single-frame 2D LiDAR, which reports where an obstacle is but not whether it is approaching. We introduce the Positional-Velocity Spatio-Temporal Attention Module (PV-STAM), a compact perception block (19,968 trainable parameters, under 3% of network capacity) that combines a per-sector scan-difference channel with a two-head self-attention mechanism and a 384→48 compression bottleneck. Policies are trained with Soft Actor-Critic across a seven-phase progressive curriculum with bidirectional demotion, scaling from static goal-seeking to fifteen simultaneously moving obstacles at 0.18 m/s, with a Hardware-Calibrated Training Mode applied in the final phase. Seven configurations were evaluated on three zero-shot benchmark arenas over 300 episodes each across three random seeds, and key control variants were validated in 130 valid physical trials on a TurtleBot3 Waffle Pi across two matched corridor scenarios. The simulation and physical evaluations give different orderings, and this dissociation is the paper’s principal result. In simulation, SAC-PV-STAM, SAC-R-PV-STAM and SAC-MLP-FS lie within 4.6 percentage points of one another on two of three benchmarks and are not separable; on hardware, they separate with large margins. SAC-R-PV-STAM reached the goal in 19 of 20 static trials and 20 of 20 trials with a moving obstacle, with no threshold violations across 40 trials, against SAC-MLP-FS (stratified p = 0.00068) and SAC-PV-STAM (stratified p = 1.1 × 10−7). A non-recurrent variant stalled with a clear floor ahead—SAC-PV-STAM commanded no forward velocity above 0.005 m/s in any of the 2746 samples of the final 91.3 s of a 122.3 s stall while the LiDAR reported 3.50 m directly ahead—and a prediction stated before the experiment, that introducing a moving obstacle would remove the stall, was confirmed (timeouts 9/10 to 0/20, p = 7 × 10−7). We further report a scan-difference contamination analysis showing up to 12.0× more spurious spikes during self-rotation on hardware, a width-matched comparison separating the effect of Huber critic loss from that of critic width, and a recurrent LSTM baseline that remains significantly inferior across all three benchmarks and unstable across seeds.
Authors
- Ayşegül Uçar (ORCID: https://orcid.org/0000-0002-5253-3779)
- Munef El Muhammed
- Anas Mahyoub Naji Saeed Alqadhi
- Mohammed Ali M. S. Bajhaw
Institutions
- University of Turku (FI)
Publication Details
- Journal
- Applied Sciences
- Published
- 2026-09-13
- DOI
- https://doi.org/10.3390/app16189083
- Primary Topic
- Advanced Neural Network Applications
- Type
- article
- Field-Weighted Citation Impact
- 0.00