LEO: Enabling Efficient Communication-Computation Pipeline for GNN via Hierarchical Caching

Training large-scale graphs with GNNs on multi-GPU platforms faces substantial feature loading overhead, leading to low resource utilization and inefficient training. Overcoming such communication bottlenecks is crucial, and current solutions achieve this by overlapping computation and communication through pipelining. However, current designs often focus only on limited operation scheduling and overlook workload imbalance within the pipeline, restricting their performance optimization in heterogeneous environments. This paper introduces LEO , a novel system to accelerate full-graph GNN training on multiple GPUs. The core of LEO lies in introducing hierarchical caching to construct a hybrid-grained software pipeline, enabling efficient computation-communication overlap both within and across GPU kernels. First, LEO introduces a hierarchical cache policy that effectively balances the workload between local and remote operations. Second, LEO tailors a cache-based pipeline and hybrid workload mapping to facilitate operation overlap. Extensive evaluations show that LEO outperforms state-of-the-art full-graph GNN systems in various settings: on averaged 7.50 ×, 11.90 ×, and 3.25 × faster than DGL, LEO-UVM, and MGG, respectively.

Authors

Institutions

Publication Details

Journal
ACM Transactions on Architecture and Code Optimization
Published
2026-09-24
DOI
https://doi.org/10.1145/3840397
Primary Topic
Graph Theory and Algorithms
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

LEO: Enabling Efficient Communication-Computation Pipeline for GNN via Hierarchical Caching

Dezun Dong, Jiaqi Si
ACM Transactions on Architecture and Code Optimization
Graph Theory and Algorithms
article

LEO: Enabling Efficient Communication-Computation Pipeline for GNN via Hierarchical Caching

Dezun Dong, Jiaqi Si
article en

Abstract

Training large-scale graphs with GNNs on multi-GPU platforms faces substantial feature loading overhead, leading to low resource utilization and inefficient training. Overcoming such communication bottlenecks is crucial, and current solutions achieve this by overlapping computation and communication through pipelining. However, current designs often focus only on limited operation scheduling and overlook workload imbalance within the pipeline, restricting their performance optimization in heterogeneous environments. This paper introduces LEO , a novel system to accelerate full-graph GNN training on multiple GPUs. The core of LEO lies in introducing hierarchical caching to construct a hybrid-grained software pipeline, enabling efficient computation-communication overlap both within and across GPU kernels. First, LEO introduces a hierarchical cache policy that effectively balances the workload between local and remote operations. Second, LEO tailors a cache-based pipeline and hybrid workload mapping to facilitate operation overlap. Extensive evaluations show that LEO outperforms state-of-the-art full-graph GNN systems in various settings: on averaged 7.50 ×, 11.90 ×, and 3.25 × faster than DGL, LEO-UVM, and MGG, respectively.

ACM Transactions on Architecture and Code Optimization
National University of Defense Technology (CN)
Decent work and economic growth
Openalex Percentile: Top 14%
Graph Theory and Algorithms
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.