Less Budget, More Memory: Diagnosing and Mitigating Non-Monotonic Activation Memory in PyTorch's Memory-Budget Partitioner
torch.compile exposes an activation_memory_budget knob that trades recomputation for memory during training. We show that the peak memory it produces is not monotone in the knob: on a 32-layer Llama-style model compiled with Inductor, a budget of 0.05 uses 2.5x the peak memory of a budget of 0.13 (904 vs. 363 MB), and a BERT model shows the same non-monotonicity below budget 0.05. We trace this to two limitations of AOTAutograd's min-cut partitioning pipeline. In the first (selection), the FLOP-maximizing knapsack that decides what to save spends small budgets on attention outputs, leaving the residual stream to be recomputed. In the second (lifetime), each recomputed value is defined once in the emitted backward graph, at its first backward use, and held until its last. Inductor's own memory estimator ranks all 14 plans we test in the same order as the measured peak (0.7% median error); the partitioner's in-tree evaluator does not. We then evaluate a post-partition, use-site rematerialization pass in the style of XLA's, which recomputes cheap values again near their late uses and leaves the saved set unchanged. At the same budget it reduces Llama's Inductor peak by 17-28% with step-time changes within run-to-run noise, removing about half of the excess, but it stays above the best stock budget. On Llama, a variant that also duplicates matrix multiplies goes 7.5-20% below the best stock budget on a dense grid of budgets, for 6-18% more step time, and makes the budget curve monotone over the budgets we tested (0.05-0.30); manual per-layer activation checkpointing reaches a lower peak still (164 MB). On BERT the excess is held dropout masks, which the pass does not regenerate because it never duplicates random operations: light rematerialization changes its peak by only 0-3%, and the all-ops variant raises it by up to 6%. Gradients of the rematerialization pass are checked against eager execution (Llama, every budget) or an unbudgeted run (BERT). We also reproduce a known bug in which recomputed dropout silently corrupts gradients under a memory budget, measure the memory cost of its proposed fix, and find that, with that fix applied, the partitioner can trigger a second, CUDA-specific random-number bug in Inductor (PyTorch issue #198333). All results come from one consumer GPU, two small models and fp32. Code and raw results: https://github.com/siddddd17/less-budget-more-memory
Authors
- Siddharth Ajith
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-05
- DOI
- https://doi.org/10.5281/zenodo.23161241
- Primary Topic
- Parallel Computing and Optimization Techniques
- Type
- preprint