LightZK: An Efficient GPU Acceleration Framework for Large-Scale Zero-Knowledge Proofs
Zero-Knowledge Proofs (ZKPs), particularly zk-SNARKs, have become a cornerstone for privacy-preserving protocols and blockchain scalability. However, the generation of proofs remains computationally prohibitive due to the immense overhead of underlying cryptographic primitives, namely Multi-Scalar Multiplication (MSM) and Fast Number Theoretic Transform (NTT). As applications scale up to billion-scale constraints, conventional GPU-accelerated frameworks tailored for smaller instances suffer from overwhelming memory usage, invalid architectural assumptions, and suboptimal end-to-end orchestration. Driven by the motivation to complete massive proof generation on a single memory-constrained node, this paper presents LightZK, an efficient GPU acceleration framework for large-scale zero-knowledge proofs based on a tile-based execution model. LightZK not only aims to support ultra-large-scale ZKP generation, but also strives to improve the practical performance of proof generation. LightZK contributes a micro-architectural diagnosis of MSM on GPUs: hardware performance counters expose the register-pressure and low instruction-level-parallelism mismatches that under-utilize the GPU during bucket-merge computation, and targeted kernel optimizations derived from this diagnosis yield a 1.5 × –2.5 × MSM speedup over the SOTA baseline. LightZK further makes large-scale NTT executable where the SOTA GPU baseline fails, via storage and structural optimizations that deliver nontrivial acceleration (2.2 × –2.5 × over a functional baseline). At the systemic level, LightZK builds on a data-dependency graph (DDG) analysis of the Groth16 proving flow to drive a three-tier pipeline orchestration: a hand-crafted DRAM-VRAM pipeline for Poly-phase that overlaps sparse matrix operations with data transfer, and a producer-consumer disk-DRAM-VRAM pipeline for Group-phase that treats each MSM tile as an independently schedulable unit. Experimental results demonstrate that LightZK achieves a 2.8 × end-to-end performance improvement over the leading executable GPU baseline at standard scales. Crucially, while existing GPU frameworks exceed available memory at billion-scale, LightZK successfully sustains 2 30 -scale proof generation with near-linear scalability on a single node.
Authors
- Xuanhua Shi (ORCID: https://orcid.org/0000-0001-8451-8656)
- Hai Jin (ORCID: https://orcid.org/0000-0002-3934-7605)
- Qiang-Sheng Hua (ORCID: https://orcid.org/0000-0002-3909-5719)
- Biran Lu (ORCID: https://orcid.org/0009-0007-1970-1340)
Institutions
- National University of Science and Technology (ZW)
Publication Details
- Journal
- ACM Transactions on Architecture and Code Optimization
- Published
- 2026-10-09
- DOI
- https://doi.org/10.1145/3857809
- Primary Topic
- Parallel Computing and Optimization Techniques
- Type
- article
- Field-Weighted Citation Impact
- 0.00