Characterising CPU-initiated High Bandwidth Memory Access

Abstract An increasing number of high-performance computers are equipped with high bandwidth memory. Their CPUs have cache-coherent access to this memory type over high-speed interconnects or as unified memory. We analyse two such systems, NVIDIA’s GH200 and AMD’s MI300A. We discuss their memory setup and measure memory access latency and throughput. We evaluate weighted memory page interleaving and strided memory access as candidates for increasing the measured throughput. Weighted interleaving utilises CPU and GPU memory simultaneously and enables exceeding the measured throughput on only CPU or GPU memory for all tested workloads. Using strided memory access, which accesses memory at fixed intervals, we achieve higher throughput than using a non-strided sequential pattern.

Authors

Institutions

Publication Details

Journal
Datenbank-Spektrum
Published
2026-10-09
DOI
https://doi.org/10.1007/s13222-026-00563-7
Primary Topic
Parallel Computing and Optimization Techniques
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Characterising CPU-initiated High Bandwidth Memory Access

Tilmann Rabl, Marcel Weisgut, Clemens Schielicke
Datenbank-Spektrum
Parallel Computing and Optimization Techniques
article

Characterising CPU-initiated High Bandwidth Memory Access

Tilmann Rabl, Marcel Weisgut, Clemens Schielicke
article en

Abstract

Abstract An increasing number of high-performance computers are equipped with high bandwidth memory. Their CPUs have cache-coherent access to this memory type over high-speed interconnects or as unified memory. We analyse two such systems, NVIDIA’s GH200 and AMD’s MI300A. We discuss their memory setup and measure memory access latency and throughput. We evaluate weighted memory page interleaving and strided memory access as candidates for increasing the measured throughput. Weighted interleaving utilises CPU and GPU memory simultaneously and enables exceeding the measured throughput on only CPU or GPU memory for all tested workloads. Using strided memory access, which accesses memory at fixed intervals, we achieve higher throughput than using a non-strided sequential pattern.

Datenbank-Spektrum
Hasso Plattner Institute (DE)
Openalex Percentile: Top 8%
Parallel Computing and Optimization Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Characterising CPU-initiated High Bandwidth Memory Access — Tilmann Rabl, Marcel Weisgut, et al. · Datenbank-Spektrum (2026) | TGRS Research Map | TGRS