Less Budget, More Memory: Diagnosing and Mitigating Non-Monotonic Activation Memory in PyTorch's Memory-Budget Partitioner

torch.compile exposes an activation_memory_budget knob that trades recomputation for memory during training. We show that the peak memory it produces is not monotone in the knob: on a 32-layer Llama-style model compiled with Inductor, a budget of 0.05 uses 2.5x the peak memory of a budget of 0.13 (904 vs. 363 MB), and a BERT model shows the same non-monotonicity below budget 0.05. We trace this to two limitations of AOTAutograd's min-cut partitioning pipeline. In the first (selection), the FLOP-maximizing knapsack that decides what to save spends small budgets on attention outputs, leaving the residual stream to be recomputed. In the second (lifetime), each recomputed value is defined once in the emitted backward graph, at its first backward use, and held until its last. Inductor's own memory estimator ranks all 14 plans we test in the same order as the measured peak (0.7% median error); the partitioner's in-tree evaluator does not. We then evaluate a post-partition, use-site rematerialization pass in the style of XLA's, which recomputes cheap values again near their late uses and leaves the saved set unchanged. At the same budget it reduces Llama's Inductor peak by 17-28% with step-time changes within run-to-run noise, removing about half of the excess, but it stays above the best stock budget. On Llama, a variant that also duplicates matrix multiplies goes 7.5-20% below the best stock budget on a dense grid of budgets, for 6-18% more step time, and makes the budget curve monotone over the budgets we tested (0.05-0.30); manual per-layer activation checkpointing reaches a lower peak still (164 MB). On BERT the excess is held dropout masks, which the pass does not regenerate because it never duplicates random operations: light rematerialization changes its peak by only 0-3%, and the all-ops variant raises it by up to 6%. Gradients of the rematerialization pass are checked against eager execution (Llama, every budget) or an unbudgeted run (BERT). We also reproduce a known bug in which recomputed dropout silently corrupts gradients under a memory budget, measure the memory cost of its proposed fix, and find that, with that fix applied, the partitioner can trigger a second, CUDA-specific random-number bug in Inductor (PyTorch issue #198333). All results come from one consumer GPU, two small models and fp32. Code and raw results: https://github.com/siddddd17/less-budget-more-memory

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-05
DOI
https://doi.org/10.5281/zenodo.23161241
Primary Topic
Parallel Computing and Optimization Techniques
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

Less Budget, More Memory: Diagnosing and Mitigating Non-Monotonic Activation Memory in PyTorch's Memory-Budget Partitioner

Siddharth Ajith
Zenodo (CERN European Organization for Nuclear Research)
Parallel Computing and Optimization Techniques
preprint

Less Budget, More Memory: Diagnosing and Mitigating Non-Monotonic Activation Memory in PyTorch's Memory-Budget Partitioner

Siddharth Ajith
preprint en

Abstract

torch.compile exposes an activation_memory_budget knob that trades recomputation for memory during training. We show that the peak memory it produces is not monotone in the knob: on a 32-layer Llama-style model compiled with Inductor, a budget of 0.05 uses 2.5x the peak memory of a budget of 0.13 (904 vs. 363 MB), and a BERT model shows the same non-monotonicity below budget 0.05. We trace this to two limitations of AOTAutograd's min-cut partitioning pipeline. In the first (selection), the FLOP-maximizing knapsack that decides what to save spends small budgets on attention outputs, leaving the residual stream to be recomputed. In the second (lifetime), each recomputed value is defined once in the emitted backward graph, at its first backward use, and held until its last. Inductor's own memory estimator ranks all 14 plans we test in the same order as the measured peak (0.7% median error); the partitioner's in-tree evaluator does not. We then evaluate a post-partition, use-site rematerialization pass in the style of XLA's, which recomputes cheap values again near their late uses and leaves the saved set unchanged. At the same budget it reduces Llama's Inductor peak by 17-28% with step-time changes within run-to-run noise, removing about half of the excess, but it stays above the best stock budget. On Llama, a variant that also duplicates matrix multiplies goes 7.5-20% below the best stock budget on a dense grid of budgets, for 6-18% more step time, and makes the budget curve monotone over the budgets we tested (0.05-0.30); manual per-layer activation checkpointing reaches a lower peak still (164 MB). On BERT the excess is held dropout masks, which the pass does not regenerate because it never duplicates random operations: light rematerialization changes its peak by only 0-3%, and the all-ops variant raises it by up to 6%. Gradients of the rematerialization pass are checked against eager execution (Llama, every budget) or an unbudgeted run (BERT). We also reproduce a known bug in which recomputed dropout silently corrupts gradients under a memory budget, measure the memory cost of its proposed fix, and find that, with that fix applied, the partitioner can trigger a second, CUDA-specific random-number bug in Inductor (PyTorch issue #198333). All results come from one consumer GPU, two small models and fp32. Code and raw results: https://github.com/siddddd17/less-budget-more-memory

Zenodo (CERN European Organization for Nuclear Research)
Parallel Computing and Optimization Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.