Optimization of Parallel Number Theoretic Transform Algorithms for Multi-Core Digital Signal Processors

The Number Theoretic Transform (NTT), a finite-field variant of FFT, is critical in cryptography and digital signal processing, but its efficiency on FT-M7032 Digital Signal Processors(DSPs) remains suboptimal due to memory bottlenecks and architectural constraints. This paper proposes a tailored optimization framework: a four-step strategy decomposes large 1D NTT into 2D transforms, outperforming traditional six-step methods by reducing memory access overhead; a VLIW/SIMD-aware microkernel with loop unrolling and register reuse boosts computation; vectorized Montgomery modular multiplication supports 16 parallel streams to address the lack of division instructions; and double buffering with pipelining overlaps computation and data transfer. Experiments show a 90.6× speedup over baselines for large-scale NTTs, with ablation studies validating each optimization. This work offers insights for NTT acceleration on DSPs and data-intensive algorithm optimization on heterogeneous processors.

Authors

Institutions

Publication Details

Journal
JUCS - Journal of Universal Computer Science
Published
2026-09-14
DOI
https://doi.org/10.3897/jucs.207845
Primary Topic
Cryptography and Residue Arithmetic
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Optimization of Parallel Number Theoretic Transform Algorithms for Multi-Core Digital Signal Processors

Dingxing Xie, Jie Liu, Peng Lin, Yang Zhang et al.
JUCS - Journal of Universal Computer Science
Cryptography and Residue Arithmetic
article

Optimization of Parallel Number Theoretic Transform Algorithms for Multi-Core Digital Signal Processors

Dingxing Xie, Jie Liu, Peng Lin, Yang Zhang, Gencheng Liu, Xiaochuan Hu, Naijun Cheng
article en

Abstract

The Number Theoretic Transform (NTT), a finite-field variant of FFT, is critical in cryptography and digital signal processing, but its efficiency on FT-M7032 Digital Signal Processors(DSPs) remains suboptimal due to memory bottlenecks and architectural constraints. This paper proposes a tailored optimization framework: a four-step strategy decomposes large 1D NTT into 2D transforms, outperforming traditional six-step methods by reducing memory access overhead; a VLIW/SIMD-aware microkernel with loop unrolling and register reuse boosts computation; vectorized Montgomery modular multiplication supports 16 parallel streams to address the lack of division instructions; and double buffering with pipelining overlaps computation and data transfer. Experiments show a 90.6× speedup over baselines for large-scale NTTs, with ablation studies validating each optimization. This work offers insights for NTT acceleration on DSPs and data-intensive algorithm optimization on heterogeneous processors.

JUCS - Journal of Universal Computer ScienceVol. 32(9)
Hunan Police Academy (CN), National University of Defense Technology (CN), Chinese PLA General Hospital (CN), Chinese People's Armed Police Force Medical College Affiliated Hospital (CN), JiangSu Armed Police General Hospital (CN), Chinese People's Armed Police General Hospital (CN)
National Natural Science Foundation of China, National University of Defense Technology, National Key Research and Development Program of China
Affordable and clean energy
Openalex Percentile: Top 4%
Cryptography and Residue Arithmetic
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.