A Controlled Comparison of GRU and LSTM Encoder–Decoder Networks for Medium-Horizon UAV Velocity-Waypoint Prediction
Accurate medium-horizon trajectory prediction, spanning roughly one to three seconds ahead, is essential for the safety and autonomy of uncrewed aerial vehicle (UAV) systems. Trajectory prediction assists UAVs in performing collision avoidance, path planning, and cooperative airspace coordination. Deep sequence models are now widely used for trajectory prediction. The gated recurrent unit (GRU) and long short-term memory (LSTM) models are among the most widely used architectures. However, published comparisons of GRU and LSTM encoder–decoder models rarely use identical data, splits, and hyperparameter budgets. This makes it difficult to attribute reported accuracy differences to the recurrent cell itself. We present a controlled comparison in which two otherwise identical velocity-based encoder–decoder networks jointly predict three future 3D velocity waypoints at +10, +20, and +30 steps ahead. Both networks share an identical 5-fold cross-validation protocol, held-out test split, fixed seed, and 81-point hyperparameter grid, for 405 runs per architecture and 810 runs total. At each architecture’s best configuration, the LSTM model reaches 7.2% lower validation Mean Squared Error (MSE) than GRU. On the held-out test set, LSTM achieves 6.5% lower test MSE under each architecture’s independently selected best configuration (best-practice comparison); a complementary matched-configuration comparison, in which each architecture is retrained under the other’s configuration, shows this advantage is concentrated in robustness to hyperparameter choice rather than in the recurrent cell alone. We also report per-waypoint, per-dimension, and Monte Carlo dropout (MC-dropout) epistemic-uncertainty metrics for both architectures. These metrics are pooled over test windows dominated by synthetic, near-planar flight data (~74%) and should not be read as general conclusions for real, free-form UAV flights. Vertical-velocity error is consistently higher than horizontal-velocity error for both models, reflecting limitations in how well the training data represent vertical movement. The MC-dropout epistemic-uncertainty intervals show 34–36% empirical coverage, substantially lower than the nominal Gaussian-reference target of 68.3% at 1σ. On our 8 × H200 GPU cluster, LSTM’s training time is 51% longer than GRU’s per run, a training-side cost that should not be read as a proxy for embedded inference cost. These results show that LSTM’s main advantage over GRU is its greater robustness to suboptimal hyperparameter settings, rather than substantially better performance when both architectures are well tuned. They also show that the uncertainty estimates from both architectures require post hoc calibration before they can be reliably used as safety bounds.
Authors
- Shokoufeh Mirzaei (ORCID: https://orcid.org/0000-0002-1102-8928)
- Shraya Ramamoorthy
- Sam Ly (ORCID: https://orcid.org/0009-0005-9166-7760)
- Siddharth Raj (ORCID: https://orcid.org/0009-0001-6179-0355)
Institutions
- Georgia Institute of Technology (US)
- California State Polytechnic University (US)
Publication Details
- Journal
- Drones
- Published
- 2026-09-25
- DOI
- https://doi.org/10.3390/drones10100732
- Primary Topic
- Air Traffic Management and Optimization
- Type
- article
- Field-Weighted Citation Impact
- 0.00