Design and FPGA Evaluation of a Multi-Operand Extension Mechanism for Tightly Coupled RISC-V Processors
Complex scalar kernels often contain instruction fragments whose computation is simple but whose operand demand exceeds the traditional two-input–one-output custom-instruction model. This operand interface bottleneck limits instruction merging in tightly coupled RISC-V acceleration. This paper proposes a multi-operand extension mechanism based on a Preload-Compute-Store architecture, which widens the operand interface over time instead of increasing the register-file read-port count. The mechanism combines U4S/D-type interface instructions, data preloading, and an output buffer to support up to six input operands and five output operands while keeping the changes to the original pipeline localized. The design was implemented in the Verilog hardware description language on Western Digital’s open-source VeeR EH1 processor core and evaluated on a Xilinx Artix-7 field-programmable gate array (FPGA) using cryptographic hashing and digital filtering kernels. Cycle-count speedups reach up to 1.15× on the hashing kernels, whereas the filtering workload gains only marginally, and part of the gain is offset by a reduction of less than 4% in the achievable clock frequency; the extension itself costs less than 2% in look-up tables and flip-flops. These results show that operand interface width is a practical constraint on tightly coupled custom-instruction acceleration, and that the mechanism is most useful for scalar fragments with sufficient operand pressure.
Authors
- Peng Lu (ORCID: https://orcid.org/0000-0001-5611-4902)
- Youping Mao
- Meijiao Yu
- Yanqing Wu
Institutions
- Chongqing University (CN)
- Chongqing Electric Power College (CN)
Publication Details
- Journal
- Electronics
- Published
- 2026-09-28
- DOI
- https://doi.org/10.3390/electronics15194464
- Primary Topic
- Parallel Computing and Optimization Techniques
- Type
- article
- Field-Weighted Citation Impact
- 0.00