ADEM: Accelerating Sparse Matrix Multiplication with Adaptive Dataflow and Efficient Merging
Sparse Matrix-Sparse Matrix Multiplication (SpMSpM) is a crucial computational kernel widely used in scientific computing and machine learning. The varying sparse patterns across different matrices pose significant challenges for conventional accelerators with fixed dataflow architectures. Although recent studies have explored dynamic dataflow approaches to better capture memory access patterns under diverse sparsity conditions, these solutions still struggle to simultaneously improve data reuse, load balance, and merging efficiency. To address these limitations, we present an SpMSpM accelerator based on adaptive dataflow and efficient merging (ADEM). ADEM is carefully designed from four key aspects. First, we propose the segmented fiber tree (SFT) data structure, which extends the conventional fiber tree through further partitioning to better support our dataflow paradigm while enhancing data reuse. We then present a cache-aware dataflow to mitigate memory overflow issues. Furthermore, ADEM decouples the multiplication and merging phases, utilizing the SFT structure to enable fine-grained task scheduling for improved load balancing. Finally, the accelerator incorporates heterogeneous merging units specifically optimized for handling two distinct types of merging operations, thereby significantly improving merger utilization. Compared with the state-of-the-art baseline system, ADEM achieves average speedups of 1.25 ×, 1.73 ×, and 1.16 × on VGG-16, ResNet-50, and SuiteSparse workloads, respectively.
Authors
- Lizhou Wu (ORCID: https://orcid.org/0000-0003-4439-7436)
- Yunping Zhao (ORCID: https://orcid.org/0000-0002-5600-3740)
- Dongsheng Li (ORCID: https://orcid.org/0000-0001-9743-2034)
- Sheng Ma (ORCID: https://orcid.org/0000-0003-1710-4060)
- Tiejun Li (ORCID: https://orcid.org/0000-0003-1509-1761)
- Shengbai Luo (ORCID: https://orcid.org/0009-0007-5551-2897)
- Bo Wang (ORCID: https://orcid.org/0009-0004-9441-0509)
- Jianmin Zhang (ORCID: https://orcid.org/0000-0002-1008-4805)
- Yuhan Tang (ORCID: https://orcid.org/0009-0003-0396-1923)
Institutions
- National Defense University (US)
- National University of Defense Technology (CN)
Publication Details
- Journal
- ACM Transactions on Architecture and Code Optimization
- Published
- 2026-09-17
- DOI
- https://doi.org/10.1145/3844726
- Primary Topic
- Parallel Computing and Optimization Techniques
- Type
- article
- Field-Weighted Citation Impact
- 0.00