Correcting Autodifferentiation in Neural ODE Training
Abstract. Does the use of autodifferentiation yield reasonable updates for deep neural networks (DNNs)? Specifically, when DNNs are designed to adhere to neural ODE architectures, can we trust the gradients provided by autodifferentiation? Through mathematical analysis and numerical evidence, we demonstrate that when neural networks employ high-order methods, such as linear multistep methods or explicit Runge–Kutta Methods (ERK), to approximate the underlying ODE flows, brute-force autodifferentiation often introduces artificial oscillations in the gradients that prevent convergence. In the case of leapfrog and 2-stage ERK, we propose simple postprocessing techniques that effectively eliminate these oscillations, correct the gradient computation, and thus return the accurate updates.
Authors
- Yewei Xu (ORCID: https://orcid.org/0000-0001-6565-058X)
- Shi Chen (ORCID: https://orcid.org/0000-0001-7864-4496)
- Qin Li (ORCID: https://orcid.org/0000-0001-9210-8948)
Institutions
- University of Wisconsin–Madison (US)
- Massachusetts Institute of Technology (US)
Publication Details
- Journal
- SIAM Journal on Applied Mathematics
- Published
- 2026-09-22
- DOI
- https://doi.org/10.1137/25m172673x
- Citations
- 1
- Primary Topic
- Model Reduction and Neural Networks
- Type
- article
- Field-Weighted Citation Impact
- 0.00
Funders
- National Science Foundation
- Office of Naval Research
- Division of Graduate Education