A Controlled Comparison of GRU and LSTM Encoder–Decoder Networks for Medium-Horizon UAV Velocity-Waypoint Prediction

Accurate medium-horizon trajectory prediction, spanning roughly one to three seconds ahead, is essential for the safety and autonomy of uncrewed aerial vehicle (UAV) systems. Trajectory prediction assists UAVs in performing collision avoidance, path planning, and cooperative airspace coordination. Deep sequence models are now widely used for trajectory prediction. The gated recurrent unit (GRU) and long short-term memory (LSTM) models are among the most widely used architectures. However, published comparisons of GRU and LSTM encoder–decoder models rarely use identical data, splits, and hyperparameter budgets. This makes it difficult to attribute reported accuracy differences to the recurrent cell itself. We present a controlled comparison in which two otherwise identical velocity-based encoder–decoder networks jointly predict three future 3D velocity waypoints at +10, +20, and +30 steps ahead. Both networks share an identical 5-fold cross-validation protocol, held-out test split, fixed seed, and 81-point hyperparameter grid, for 405 runs per architecture and 810 runs total. At each architecture’s best configuration, the LSTM model reaches 7.2% lower validation Mean Squared Error (MSE) than GRU. On the held-out test set, LSTM achieves 6.5% lower test MSE under each architecture’s independently selected best configuration (best-practice comparison); a complementary matched-configuration comparison, in which each architecture is retrained under the other’s configuration, shows this advantage is concentrated in robustness to hyperparameter choice rather than in the recurrent cell alone. We also report per-waypoint, per-dimension, and Monte Carlo dropout (MC-dropout) epistemic-uncertainty metrics for both architectures. These metrics are pooled over test windows dominated by synthetic, near-planar flight data (~74%) and should not be read as general conclusions for real, free-form UAV flights. Vertical-velocity error is consistently higher than horizontal-velocity error for both models, reflecting limitations in how well the training data represent vertical movement. The MC-dropout epistemic-uncertainty intervals show 34–36% empirical coverage, substantially lower than the nominal Gaussian-reference target of 68.3% at 1σ. On our 8 × H200 GPU cluster, LSTM’s training time is 51% longer than GRU’s per run, a training-side cost that should not be read as a proxy for embedded inference cost. These results show that LSTM’s main advantage over GRU is its greater robustness to suboptimal hyperparameter settings, rather than substantially better performance when both architectures are well tuned. They also show that the uncertainty estimates from both architectures require post hoc calibration before they can be reliably used as safety bounds.

Authors

Institutions

Publication Details

Journal
Drones
Published
2026-09-25
DOI
https://doi.org/10.3390/drones10100732
Primary Topic
Air Traffic Management and Optimization
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

A Controlled Comparison of GRU and LSTM Encoder–Decoder Networks for Medium-Horizon UAV Velocity-Waypoint Prediction

Shokoufeh Mirzaei, Shraya Ramamoorthy, Sam Ly, Siddharth Raj
Drones
Air Traffic Management and Optimization
article

A Controlled Comparison of GRU and LSTM Encoder–Decoder Networks for Medium-Horizon UAV Velocity-Waypoint Prediction

Shokoufeh Mirzaei, Shraya Ramamoorthy, Sam Ly, Siddharth Raj
article en

Abstract

Accurate medium-horizon trajectory prediction, spanning roughly one to three seconds ahead, is essential for the safety and autonomy of uncrewed aerial vehicle (UAV) systems. Trajectory prediction assists UAVs in performing collision avoidance, path planning, and cooperative airspace coordination. Deep sequence models are now widely used for trajectory prediction. The gated recurrent unit (GRU) and long short-term memory (LSTM) models are among the most widely used architectures. However, published comparisons of GRU and LSTM encoder–decoder models rarely use identical data, splits, and hyperparameter budgets. This makes it difficult to attribute reported accuracy differences to the recurrent cell itself. We present a controlled comparison in which two otherwise identical velocity-based encoder–decoder networks jointly predict three future 3D velocity waypoints at +10, +20, and +30 steps ahead. Both networks share an identical 5-fold cross-validation protocol, held-out test split, fixed seed, and 81-point hyperparameter grid, for 405 runs per architecture and 810 runs total. At each architecture’s best configuration, the LSTM model reaches 7.2% lower validation Mean Squared Error (MSE) than GRU. On the held-out test set, LSTM achieves 6.5% lower test MSE under each architecture’s independently selected best configuration (best-practice comparison); a complementary matched-configuration comparison, in which each architecture is retrained under the other’s configuration, shows this advantage is concentrated in robustness to hyperparameter choice rather than in the recurrent cell alone. We also report per-waypoint, per-dimension, and Monte Carlo dropout (MC-dropout) epistemic-uncertainty metrics for both architectures. These metrics are pooled over test windows dominated by synthetic, near-planar flight data (~74%) and should not be read as general conclusions for real, free-form UAV flights. Vertical-velocity error is consistently higher than horizontal-velocity error for both models, reflecting limitations in how well the training data represent vertical movement. The MC-dropout epistemic-uncertainty intervals show 34–36% empirical coverage, substantially lower than the nominal Gaussian-reference target of 68.3% at 1σ. On our 8 × H200 GPU cluster, LSTM’s training time is 51% longer than GRU’s per run, a training-side cost that should not be read as a proxy for embedded inference cost. These results show that LSTM’s main advantage over GRU is its greater robustness to suboptimal hyperparameter settings, rather than substantially better performance when both architectures are well tuned. They also show that the uncertainty estimates from both architectures require post hoc calibration before they can be reliably used as safety bounds.

DronesVol. 10(10)
Georgia Institute of Technology (US), California State Polytechnic University (US)
Openalex Percentile: Top 8%
Air Traffic Management and Optimization
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.