Optimizing Pipeline Parallelism for Deep Learning With Activation Checkpointing

ABSTRACT This paper accelerates fast‐forward model training that applies pipeline parallelism and activation checkpointing implemented by PyTorch. We prove that it is NP‐hard to select checkpoints that minimize the training time under a given memory limit and PyTorch resource usage. We design a dynamic programming algorithm to partition a model into stages that is mathematically proven to minimize the maximum training time among stages. We design a dynamic programming algorithm to select checkpoints that are mathematically proven to minimize the stage training time under a memory limit and PyTorch resource usage. We design an FPTAS for checkpoint selection by rounding the training time of layers. Empirically, our algorithms achieve speedups of up to 2.07 over the state‐of‐the‐art algorithms across models of various sizes and structures. The empirical stage training time corroborates that our algorithms achieve a more balanced stage training time across stages, and our checkpoints have a shorter stage training time for a single stage than the state‐of‐the‐art algorithm.

Authors

Institutions

Publication Details

Journal
Concurrency and Computation Practice and Experience
Published
2026-09-29
DOI
https://doi.org/10.1002/cpe.70975
Primary Topic
Distributed systems and fault tolerance
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Optimizing Pipeline Parallelism for Deep Learning With Activation Checkpointing

Pangfeng Liu, Ding‐Yong Hong, Jan‐Jan Wu, Ming-Yen Chiang et al.
Concurrency and Computation Practice and Experience
Distributed systems and fault tolerance
article

Optimizing Pipeline Parallelism for Deep Learning With Activation Checkpointing

Pangfeng Liu, Ding‐Yong Hong, Jan‐Jan Wu, Ming-Yen Chiang, Tzu‐Hsien Tsai
article en

Abstract

ABSTRACT This paper accelerates fast‐forward model training that applies pipeline parallelism and activation checkpointing implemented by PyTorch. We prove that it is NP‐hard to select checkpoints that minimize the training time under a given memory limit and PyTorch resource usage. We design a dynamic programming algorithm to partition a model into stages that is mathematically proven to minimize the maximum training time among stages. We design a dynamic programming algorithm to select checkpoints that are mathematically proven to minimize the stage training time under a memory limit and PyTorch resource usage. We design an FPTAS for checkpoint selection by rounding the training time of layers. Empirically, our algorithms achieve speedups of up to 2.07 over the state‐of‐the‐art algorithms across models of various sizes and structures. The empirical stage training time corroborates that our algorithms achieve a more balanced stage training time across stages, and our checkpoints have a shorter stage training time for a single stage than the state‐of‐the‐art algorithm.

Concurrency and Computation Practice and ExperienceVol. 38(19)
National Taiwan University (TW), Institute of Information Science, Academia Sinica (TW)
Openalex Percentile: Top 9%
Distributed systems and fault tolerance
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Optimizing Pipeline Parallelism for Deep Learning With Activation Checkpointing — Pangfeng Liu, Ding‐Yong Hong, et al. · Concurrency and Computation Practice and Experience (2026) | TGRS Research Map | TGRS