Characterising CPU-initiated High Bandwidth Memory Access
Abstract An increasing number of high-performance computers are equipped with high bandwidth memory. Their CPUs have cache-coherent access to this memory type over high-speed interconnects or as unified memory. We analyse two such systems, NVIDIA’s GH200 and AMD’s MI300A. We discuss their memory setup and measure memory access latency and throughput. We evaluate weighted memory page interleaving and strided memory access as candidates for increasing the measured throughput. Weighted interleaving utilises CPU and GPU memory simultaneously and enables exceeding the measured throughput on only CPU or GPU memory for all tested workloads. Using strided memory access, which accesses memory at fixed intervals, we achieve higher throughput than using a non-strided sequential pattern.
Authors
- Tilmann Rabl (ORCID: https://orcid.org/0009-0009-3335-8045)
- Marcel Weisgut (ORCID: https://orcid.org/0009-0002-8973-6403)
- Clemens Schielicke (ORCID: https://orcid.org/0009-0007-6844-6286)
Institutions
- Hasso Plattner Institute (DE)
Publication Details
- Journal
- Datenbank-Spektrum
- Published
- 2026-10-09
- DOI
- https://doi.org/10.1007/s13222-026-00563-7
- Primary Topic
- Parallel Computing and Optimization Techniques
- Type
- article
- Field-Weighted Citation Impact
- 0.00