OLA+: Multi-FPGA O ver l ay A ccelerator System for Fully Homomorphic Encryption

Fully Homomorphic Encryption (FHE) is a promising technique for privacy preserving computation. However, homomorphically encrypted operations are orders of magnitude slower than the corresponding unencrypted operations due to high computation and memory bandwidth requirements. FPGAs are attractive platforms for accelerating FHE workloads, but manually programming FPGAs is challenging, as different FHE parameters and operations require different mapping strategies. We propose OLA+, a scalable overlay accelerator system for FHE, designed to run efficiently on multi-FPGA platforms. OLA+ eliminates the need for manual and time-consuming FPGA programming by providing a Python-based interface and a compiler that efficiently maps OLA+ programs to hardware instructions for execution. The hardware architecture and instruction set are co-designed to accelerate common FHE primitives, while the compiler maps all FHE operations to these primitives for efficient execution. We propose a compile-time partitioning strategy that parallelizes computation both across multiple ciphertexts and within individual ciphertexts, enabling efficient execution of OLA+ programs on multiple FPGA accelerators. OLA+ features two optimizations to address memory bandwidth challenges in FHE computation. First, we propose an asynchronous dataflow execution model where the compiler manages data processing order and the architecture enforces it at run time. This approach enables guaranteed data reuse via on-chip SRAMs. Second, we design latency-aware instruction scheduling in the compiler to reduce data reuse distance and overlap data transfers with computation. We implement the overlay accelerator on AMD Alveo U280 FPGAs. We evaluate the effectiveness of the proposed overlay accelerator across a range of FHE parameters by executing various FHE operations, FHE linear algebra benchmarks, and end-to-end ML training and inference. Experimental results show that running end-to-end HE ML training and inference tasks on OLA+ with 8 overlay accelerators achieves speedups of up to \(8835\times\) and \(14.6\times\) over state-of-the-art CPU and GPU implementations, respectively.

Authors

Institutions

Publication Details

Journal
ACM Transactions on Reconfigurable Technology and Systems
Published
2026-10-06
DOI
https://doi.org/10.1145/3856805
Primary Topic
Embedded Systems Design Techniques
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

OLA+: Multi-FPGA O ver l ay A ccelerator System for Fully Homomorphic Encryption

Viktor K. Prasanna, Rajgopal Kannan, Yang Yang
ACM Transactions on Reconfigurable Technology and Systems
Embedded Systems Design Techniques
article

OLA+: Multi-FPGA O ver l ay A ccelerator System for Fully Homomorphic Encryption

Viktor K. Prasanna, Rajgopal Kannan, Yang Yang
article en

Abstract

Fully Homomorphic Encryption (FHE) is a promising technique for privacy preserving computation. However, homomorphically encrypted operations are orders of magnitude slower than the corresponding unencrypted operations due to high computation and memory bandwidth requirements. FPGAs are attractive platforms for accelerating FHE workloads, but manually programming FPGAs is challenging, as different FHE parameters and operations require different mapping strategies. We propose OLA+, a scalable overlay accelerator system for FHE, designed to run efficiently on multi-FPGA platforms. OLA+ eliminates the need for manual and time-consuming FPGA programming by providing a Python-based interface and a compiler that efficiently maps OLA+ programs to hardware instructions for execution. The hardware architecture and instruction set are co-designed to accelerate common FHE primitives, while the compiler maps all FHE operations to these primitives for efficient execution. We propose a compile-time partitioning strategy that parallelizes computation both across multiple ciphertexts and within individual ciphertexts, enabling efficient execution of OLA+ programs on multiple FPGA accelerators. OLA+ features two optimizations to address memory bandwidth challenges in FHE computation. First, we propose an asynchronous dataflow execution model where the compiler manages data processing order and the architecture enforces it at run time. This approach enables guaranteed data reuse via on-chip SRAMs. Second, we design latency-aware instruction scheduling in the compiler to reduce data reuse distance and overlap data transfers with computation. We implement the overlay accelerator on AMD Alveo U280 FPGAs. We evaluate the effectiveness of the proposed overlay accelerator across a range of FHE parameters by executing various FHE operations, FHE linear algebra benchmarks, and end-to-end ML training and inference. Experimental results show that running end-to-end HE ML training and inference tasks on OLA+ with 8 overlay accelerators achieves speedups of up to \(8835\times\) and \(14.6\times\) over state-of-the-art CPU and GPU implementations, respectively.

ACM Transactions on Reconfigurable Technology and Systems
University of Southern California (US), United States Department of the Army (US), DEVCOM Army Research Laboratory (US), United States Army (US), United States Army Research Office (US)
Openalex Percentile: Top 7%
Embedded Systems Design Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.