On the Performance of Stochastic Gradient Methods with Momentum in Time-Varying Regimes
Abstract We explore how Stochastic Gradient methods with momentum perform within a time-varying framework by establishing bounds on their tracking errors, specifically focusing on quadratic cases. Notably, we find that momentum methods achieve, in high drift-to-noise regimes , i.e., when the rate of change of the dynamic optimum prevails on the variance of the gradient noise, smaller neighborhood of convergence compared to stochastic gradient descent. To the best of our knowledge, this is the first proof that a momentum method can improve upon stochastic gradient descent’s tracking error bounds in a time-varying setting. We then investigate, for a given learning rate, the optimal choice of the momentum parameter that minimizes the tracking error bound.
Authors
- Enrico Bernardi (ORCID: https://orcid.org/0000-0003-3923-1407)
- Christopher S. A. Lauria (ORCID: https://orcid.org/0009-0004-3554-0215)
- Alberto Lanconelli (ORCID: https://orcid.org/0000-0001-8248-2151)
Institutions
- University of Bologna (IT)
Publication Details
- Journal
- Applied Mathematics & Optimization
- Published
- 2026-09-18
- DOI
- https://doi.org/10.1007/s00245-026-10517-w
- Primary Topic
- Stochastic Gradient Optimization Techniques
- Type
- article
- Field-Weighted Citation Impact
- 0.00