Convergence Dynamics and Implicit Regularization in Gradient-Based Optimization: A Study Using Computer Experiments
Abstract: Machine learning models learn by changing their parameters to make a loss function smaller, and the most common way to do this is gradient descent. This paper studies four questions about gradient descent with a mix of mathematics and computer experiments written in Python. First, how do the learning rate and the shape of the loss function control whether and how fast gradient descent converges? Second, when many different solutions fit the training data perfectly, which one does gradient descent choose? On a simple quadratic loss, simulations matched the formula that the learning rate must stay below 2/a, and the number of steps needed grew with the condition number κ roughly like κ ln(1/ε). With 10 data points and 50 weights, gradient descent started at zero found the minimum-norm solution (distance 4.2e-15), which was much smaller (norm 3.57) than the typical other perfect solution (average norm 7.26). Finally, on noisy data, stopping gradient descent early reduced the test error from 12.49 to 5.60 in one example, and behaved much like ridge regression. These results use simple linear models, so they illustrate ideas from deep learning research but do not prove them for neural networks. Keywords: gradient descent, learning rate, condition number, implicit regularisation, early stopping
Authors
- Mrunmayee Kulkarni
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-05
- DOI
- https://doi.org/10.5281/zenodo.23149288
- Primary Topic
- Stochastic Gradient Optimization Techniques
- Type
- article
- Field-Weighted Citation Impact
- 0.00