Optimizing Pipeline Parallelism for Deep Learning With Activation Checkpointing
ABSTRACT This paper accelerates fast‐forward model training that applies pipeline parallelism and activation checkpointing implemented by PyTorch. We prove that it is NP‐hard to select checkpoints that minimize the training time under a given memory limit and PyTorch resource usage. We design a dynamic programming algorithm to partition a model into stages that is mathematically proven to minimize the maximum training time among stages. We design a dynamic programming algorithm to select checkpoints that are mathematically proven to minimize the stage training time under a memory limit and PyTorch resource usage. We design an FPTAS for checkpoint selection by rounding the training time of layers. Empirically, our algorithms achieve speedups of up to 2.07 over the state‐of‐the‐art algorithms across models of various sizes and structures. The empirical stage training time corroborates that our algorithms achieve a more balanced stage training time across stages, and our checkpoints have a shorter stage training time for a single stage than the state‐of‐the‐art algorithm.
Authors
- Pangfeng Liu (ORCID: https://orcid.org/0000-0002-5466-9960)
- Ding‐Yong Hong (ORCID: https://orcid.org/0000-0002-7649-7581)
- Jan‐Jan Wu (ORCID: https://orcid.org/0000-0003-1722-4361)
- Ming-Yen Chiang
- Tzu‐Hsien Tsai (ORCID: https://orcid.org/0009-0008-8132-3122)
Institutions
- National Taiwan University (TW)
- Institute of Information Science, Academia Sinica (TW)
Publication Details
- Journal
- Concurrency and Computation Practice and Experience
- Published
- 2026-09-29
- DOI
- https://doi.org/10.1002/cpe.70975
- Primary Topic
- Distributed systems and fault tolerance
- Type
- article
- Field-Weighted Citation Impact
- 0.00