Dissecting the NVIDIA Hopper Architecture through Microbenchmarking and Multiple Level Analysis

This study presents a comprehensive multi-level analysis of the NVIDIA Hopper GPU architecture, focusing on its performance characteristics and novel features. We benchmark Hopper’s memory subsystem, highlighting improvements in the L2 partitioned cache and global memory access compared to Ampere and Ada Lovelace. The evaluation of Hopper’s fourth-generation tensor cores reveals the benefits of FP8 precision and asynchronous wgmma instructions for matrix operations. Additionally, we investigate the performance of DPX instructions for dynamic programming, distributed shared memory (DSM) for inter-SM communication, and the Tensor Memory Accelerator (TMA) for asynchronous data movement. Through multi-level evaluation, we discover that the Hopper architecture demonstrates significant acceleration potential in real-world applications. For instance, the asynchronous programming model supported by TMA achieves a 1.5× speedup in matrix multiplication, FP8 delivers nearly double the performance of FP16, and DPX instructions accelerate a computational biology algorithm by at least 4.75×. Our findings provide actionable insights for optimizing compute-intensive workloads, from AI training to bioinformatics, on Hopper GPUs.

Authors

Institutions

Publication Details

Journal
ACM Transactions on Computer Systems
Published
2026-10-05
DOI
https://doi.org/10.1145/3856813
Citations
1
Primary Topic
Parallel Computing and Optimization Techniques
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Dissecting the NVIDIA Hopper Architecture through Microbenchmarking and Multiple Level Analysis

Ruibo Fan, Hongyuan Liu, Zeyu Li, Xiaowen Chu et al.
1 citations
ACM Transactions on Computer Systems
Parallel Computing and Optimization Techniques
article

Dissecting the NVIDIA Hopper Architecture through Microbenchmarking and Multiple Level Analysis

Ruibo Fan, Hongyuan Liu, Zeyu Li, Xiaowen Chu, Qiang Wang, Weile Luo, Dayou Du
article en
1 citations

Abstract

This study presents a comprehensive multi-level analysis of the NVIDIA Hopper GPU architecture, focusing on its performance characteristics and novel features. We benchmark Hopper’s memory subsystem, highlighting improvements in the L2 partitioned cache and global memory access compared to Ampere and Ada Lovelace. The evaluation of Hopper’s fourth-generation tensor cores reveals the benefits of FP8 precision and asynchronous wgmma instructions for matrix operations. Additionally, we investigate the performance of DPX instructions for dynamic programming, distributed shared memory (DSM) for inter-SM communication, and the Tensor Memory Accelerator (TMA) for asynchronous data movement. Through multi-level evaluation, we discover that the Hopper architecture demonstrates significant acceleration potential in real-world applications. For instance, the asynchronous programming model supported by TMA achieves a 1.5× speedup in matrix multiplication, FP8 delivers nearly double the performance of FP16, and DPX instructions accelerate a computational biology algorithm by at least 4.75×. Our findings provide actionable insights for optimizing compute-intensive workloads, from AI training to bioinformatics, on Hopper GPUs.

ACM Transactions on Computer Systems
Stevens Institute of Technology (US), Hong Kong University of Science and Technology (HK), Harbin Institute of Technology (CN), Guangzhou HKUST Fok Ying Tung Research Institute (CN), The Hong Kong University of Science and Technology (Guangzhou) (CN)
National Natural Science Foundation of China
Openalex Percentile: Top 100%
Parallel Computing and Optimization Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.