RecDM: Efficient Training System for Large-Scale Recommendation Models on Disaggregated Memory

The embedding tables in Deep Learning Recommendation Models (DLRMs) require significant memory capacity and bandwidth but relatively lower computing power, making it economically inefficient to scale by adding more GPUs solely to meet memory requirements. Recent advances in Compute Express Link (CXL) and near-data processing (NDP) offer a promising path for expanding system memory to support DLRM training with large embedding tables. However, existing CXL+NDP designs are limited to single-GPU scenarios and fail to address key challenges in multi-GPU settings, like memory contention and embedding placement. To address this issue, we propose RecDM , an efficient training system for large-scale Rec ommendation models on D isaggregated M emory. RecDM features a modular, many-to-many CXL-based architecture with lightweight NDP units integrated into the CXL controller of each memory expansion unit, avoiding changes to DRAM chips or DIMM organization. To optimize training throughput, RecDM introduces a hierarchical memory device allocation strategy that balances memory bandwidth and capacity by combining shared and exclusive device mappings. Furthermore, RecDM proposes a bandwidth-driven 2D embedding table sharding method, enabling fine-grained placement across heterogeneous memory hierarchies and devices. Finally, RecDM incorporates an input-adaptive communication routing mechanism combined with a pipelined GPU-CXL execution model to reduce synchronization and data movement overhead. Comprehensive experiments demonstrate that RecDM achieves an average 10.2 × speedup over prior work that offloads embedding tables to host DRAM, and outperforms state-of-the-art CXL+NDP solution by an average of 2.2 ×.

Authors

Institutions

Publication Details

Journal
ACM Transactions on Architecture and Code Optimization
Published
2026-09-28
DOI
https://doi.org/10.1145/3848634
Primary Topic
Advanced Neural Network Applications
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

RecDM: Efficient Training System for Large-Scale Recommendation Models on Disaggregated Memory

Yufei Ding, Zhongkai Yu, Yangwook Kang, Yuke Wang et al.
ACM Transactions on Architecture and Code Optimization
Advanced Neural Network Applications
article

RecDM: Efficient Training System for Large-Scale Recommendation Models on Disaggregated Memory

Yufei Ding, Zhongkai Yu, Yangwook Kang, Yuke Wang, Xulong Tang, Liu Liu, Yikai Li, Zheng Wang, Yichen Lin, Kaijian Wang
article en

Abstract

The embedding tables in Deep Learning Recommendation Models (DLRMs) require significant memory capacity and bandwidth but relatively lower computing power, making it economically inefficient to scale by adding more GPUs solely to meet memory requirements. Recent advances in Compute Express Link (CXL) and near-data processing (NDP) offer a promising path for expanding system memory to support DLRM training with large embedding tables. However, existing CXL+NDP designs are limited to single-GPU scenarios and fail to address key challenges in multi-GPU settings, like memory contention and embedding placement. To address this issue, we propose RecDM , an efficient training system for large-scale Rec ommendation models on D isaggregated M emory. RecDM features a modular, many-to-many CXL-based architecture with lightweight NDP units integrated into the CXL controller of each memory expansion unit, avoiding changes to DRAM chips or DIMM organization. To optimize training throughput, RecDM introduces a hierarchical memory device allocation strategy that balances memory bandwidth and capacity by combining shared and exclusive device mappings. Furthermore, RecDM proposes a bandwidth-driven 2D embedding table sharding method, enabling fine-grained placement across heterogeneous memory hierarchies and devices. Finally, RecDM incorporates an input-adaptive communication routing mechanism combined with a pipelined GPU-CXL execution model to reduce synchronization and data movement overhead. Comprehensive experiments demonstrate that RecDM achieves an average 10.2 × speedup over prior work that offloads embedding tables to host DRAM, and outperforms state-of-the-art CXL+NDP solution by an average of 2.2 ×.

ACM Transactions on Architecture and Code Optimization
Rensselaer Polytechnic Institute (US), University of Pittsburgh (US), University of California San Diego (US), Samsung Electronics (South Korea) (KR), Rice University (US)
Openalex Percentile: Top 14%
Advanced Neural Network Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.