Demystifying SpMV Performance: Bottleneck Characterization and Open Challenges
Sparse Matrix-Vector Multiplication (SpMV) is a critical computational kernel in scientific computing, graph processing, and machine learning. Despite its importance, SpMV achieves only a small fraction of peak performance, varying substantially across matrices and architectures. This paper comprehensively analyzes SpMV performance on modern CPUs and GPUs, identifying key bottlenecks: memory bandwidth limitations, low instruction-level parallelism (ILP), load imbalance, and memory latency overheads. We characterize matrices via four structural features, each linked to a bottleneck, to enable structured performance analysis. We also develop and publicly release a matrix generator producing artificial matrices across a broad range of structural properties, enabling extensive benchmarking from high-end HPC accelerators to consumer-grade GPUs. While GPUs achieve superior SpMV performance, modern CPUs with larger caches and increased core counts remain competitive, particularly for small-to-medium sized matrices. Evaluating numerous storage formats, we demonstrate the efficiency of vendor-customized implementations on GPUs and the potential of research formats to address specific CPU performance challenges. Expanding to iterative solvers (BiCG) and FP32 arithmetic reveals non-obvious architectural asymmetries: GPU throughput is resilient to solver-level cache pressure while CPUs degrade significantly, and FP32 scaling yields architecture-specific gains, from bandwidth improvements to unlocked compute capacity on consumer GPUs. Finally, we identify persistent open challenges: memory bandwidth constraints, irregular matrix structures, and short unbalanced rows, underscoring the need for further optimization and architectural improvements.
Authors
- Nectarios Koziris (ORCID: https://orcid.org/0000-0002-4890-8427)
- Georgios Goumas (ORCID: https://orcid.org/0000-0001-7811-4831)
- Panagiotis Mpakos (ORCID: https://orcid.org/0009-0006-5148-6903)
- Νικέλα Παπαδοπούλου (ORCID: https://orcid.org/0000-0003-2141-5654)
- Dimitrios Galanopoulos (ORCID: https://orcid.org/0009-0007-2257-8670)
- Petros Anastasiadis (ORCID: https://orcid.org/0000-0001-7821-3610)
- Ioanna Tasou
Institutions
- National Technical University of Athens (GR)
- University of Glasgow (GB)
Publication Details
- Journal
- ACM Transactions on Architecture and Code Optimization
- Published
- 2026-10-05
- DOI
- https://doi.org/10.1145/3856990
- Primary Topic
- Parallel Computing and Optimization Techniques
- Type
- article
- Field-Weighted Citation Impact
- 0.00