SCALABLE FPGA-BASED DEEP LEARNING ACCELERATOR USING TILED MATRIX MULTIPLICATION AND PIPELINED PROCESSING

Deep neural networks provide strong performance in recognition, prediction, and representation-learning tasks, but their computational intensity and memory-access requirements make efficient deployment difficult on conventional processors. This paper presents a publication-oriented formulation of a scalable Deep Learning Accelerator Unit (DLAU) implemented on a field-programmable gate array (FPGA). The architecture targets the dominant kernels of deep learning workloads by combining tiled data processing, explicit on-chip buffering, streaming communication, and a three-stage fully pipelined datapath. The processing chain consists of a Tiled Matrix Multiplication Unit (TMMU), a Part Sum Accumulation Unit (PSAU), and an Activation Function Acceleration Unit (AFAU). Tiling decouples the supported neural-network size from the fixed amount of FPGA arithmetic resources, while FIFO buffering and pipelined execution allow the three units to operate concurrently. Profiling data in the source design shows that matrix multiplication accounts for 98.6%, 98.2%, and 99.1% of the runtime of feedforward, restricted Boltzmann machine, and back-propagation operations, respectively, motivating the emphasis on a reusable matrix-processing engine. The available FPGA prototype results report a speedup of up to 36.1× over an Intel Core2 processor with a power consumption of approximately 234 mW. The resulting architecture demonstrates how data locality, resource reuse, and coarse-grained pipelining can be combined to build a flexible accelerator for large-scale deep-learning workloads under the area, memory, and energy constraints of reconfigurable hardware.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-14
DOI
https://doi.org/10.5281/zenodo.22754267
Primary Topic
Advanced Neural Network Applications
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

SCALABLE FPGA-BASED DEEP LEARNING ACCELERATOR USING TILED MATRIX MULTIPLICATION AND PIPELINED PROCESSING

Guguloth Laxman, N. Bhavana, M. Shobha Rani
Zenodo (CERN European Organization for Nuclear Research)
Advanced Neural Network Applications
article

SCALABLE FPGA-BASED DEEP LEARNING ACCELERATOR USING TILED MATRIX MULTIPLICATION AND PIPELINED PROCESSING

Guguloth Laxman, N. Bhavana, M. Shobha Rani
article en

Abstract

Deep neural networks provide strong performance in recognition, prediction, and representation-learning tasks, but their computational intensity and memory-access requirements make efficient deployment difficult on conventional processors. This paper presents a publication-oriented formulation of a scalable Deep Learning Accelerator Unit (DLAU) implemented on a field-programmable gate array (FPGA). The architecture targets the dominant kernels of deep learning workloads by combining tiled data processing, explicit on-chip buffering, streaming communication, and a three-stage fully pipelined datapath. The processing chain consists of a Tiled Matrix Multiplication Unit (TMMU), a Part Sum Accumulation Unit (PSAU), and an Activation Function Acceleration Unit (AFAU). Tiling decouples the supported neural-network size from the fixed amount of FPGA arithmetic resources, while FIFO buffering and pipelined execution allow the three units to operate concurrently. Profiling data in the source design shows that matrix multiplication accounts for 98.6%, 98.2%, and 99.1% of the runtime of feedforward, restricted Boltzmann machine, and back-propagation operations, respectively, motivating the emphasis on a reusable matrix-processing engine. The available FPGA prototype results report a speedup of up to 36.1× over an Intel Core2 processor with a power consumption of approximately 234 mW. The resulting architecture demonstrates how data locality, resource reuse, and coarse-grained pipelining can be combined to build a flexible accelerator for large-scale deep-learning workloads under the area, memory, and energy constraints of reconfigurable hardware.

Zenodo (CERN European Organization for Nuclear Research)
Grammar School (SK)
Openalex Percentile: Top 13%
Advanced Neural Network Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

SCALABLE FPGA-BASED DEEP LEARNING ACCELERATOR USING TILED MATRIX MULTIPLICATION AND PIPELINED PROCESSING — Guguloth Laxman, N. Bhavana, et al. · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS