OLA+: Multi-FPGA O ver l ay A ccelerator System for Fully Homomorphic Encryption
Fully Homomorphic Encryption (FHE) is a promising technique for privacy preserving computation. However, homomorphically encrypted operations are orders of magnitude slower than the corresponding unencrypted operations due to high computation and memory bandwidth requirements. FPGAs are attractive platforms for accelerating FHE workloads, but manually programming FPGAs is challenging, as different FHE parameters and operations require different mapping strategies. We propose OLA+, a scalable overlay accelerator system for FHE, designed to run efficiently on multi-FPGA platforms. OLA+ eliminates the need for manual and time-consuming FPGA programming by providing a Python-based interface and a compiler that efficiently maps OLA+ programs to hardware instructions for execution. The hardware architecture and instruction set are co-designed to accelerate common FHE primitives, while the compiler maps all FHE operations to these primitives for efficient execution. We propose a compile-time partitioning strategy that parallelizes computation both across multiple ciphertexts and within individual ciphertexts, enabling efficient execution of OLA+ programs on multiple FPGA accelerators. OLA+ features two optimizations to address memory bandwidth challenges in FHE computation. First, we propose an asynchronous dataflow execution model where the compiler manages data processing order and the architecture enforces it at run time. This approach enables guaranteed data reuse via on-chip SRAMs. Second, we design latency-aware instruction scheduling in the compiler to reduce data reuse distance and overlap data transfers with computation. We implement the overlay accelerator on AMD Alveo U280 FPGAs. We evaluate the effectiveness of the proposed overlay accelerator across a range of FHE parameters by executing various FHE operations, FHE linear algebra benchmarks, and end-to-end ML training and inference. Experimental results show that running end-to-end HE ML training and inference tasks on OLA+ with 8 overlay accelerators achieves speedups of up to \(8835\times\) and \(14.6\times\) over state-of-the-art CPU and GPU implementations, respectively.
Authors
- Viktor K. Prasanna (ORCID: https://orcid.org/0000-0002-1609-8589)
- Rajgopal Kannan (ORCID: https://orcid.org/0000-0001-8736-3012)
- Yang Yang (ORCID: https://orcid.org/0000-0002-7718-1925)
Institutions
- University of Southern California (US)
- United States Department of the Army (US)
- DEVCOM Army Research Laboratory (US)
- United States Army (US)
- United States Army Research Office (US)
Publication Details
- Journal
- ACM Transactions on Reconfigurable Technology and Systems
- Published
- 2026-10-06
- DOI
- https://doi.org/10.1145/3856805
- Primary Topic
- Embedded Systems Design Techniques
- Type
- article
- Field-Weighted Citation Impact
- 0.00