LEO: Enabling Efficient Communication-Computation Pipeline for GNN via Hierarchical Caching
Training large-scale graphs with GNNs on multi-GPU platforms faces substantial feature loading overhead, leading to low resource utilization and inefficient training. Overcoming such communication bottlenecks is crucial, and current solutions achieve this by overlapping computation and communication through pipelining. However, current designs often focus only on limited operation scheduling and overlook workload imbalance within the pipeline, restricting their performance optimization in heterogeneous environments. This paper introduces LEO , a novel system to accelerate full-graph GNN training on multiple GPUs. The core of LEO lies in introducing hierarchical caching to construct a hybrid-grained software pipeline, enabling efficient computation-communication overlap both within and across GPU kernels. First, LEO introduces a hierarchical cache policy that effectively balances the workload between local and remote operations. Second, LEO tailors a cache-based pipeline and hybrid workload mapping to facilitate operation overlap. Extensive evaluations show that LEO outperforms state-of-the-art full-graph GNN systems in various settings: on averaged 7.50 ×, 11.90 ×, and 3.25 × faster than DGL, LEO-UVM, and MGG, respectively.
Authors
- Dezun Dong (ORCID: https://orcid.org/0000-0001-6243-8479)
- Jiaqi Si (ORCID: https://orcid.org/0000-0002-9959-611X)
Institutions
- National University of Defense Technology (CN)
Publication Details
- Journal
- ACM Transactions on Architecture and Code Optimization
- Published
- 2026-09-24
- DOI
- https://doi.org/10.1145/3840397
- Primary Topic
- Graph Theory and Algorithms
- Type
- article
- Field-Weighted Citation Impact
- 0.00