RecDM: Efficient Training System for Large-Scale Recommendation Models on Disaggregated Memory
The embedding tables in Deep Learning Recommendation Models (DLRMs) require significant memory capacity and bandwidth but relatively lower computing power, making it economically inefficient to scale by adding more GPUs solely to meet memory requirements. Recent advances in Compute Express Link (CXL) and near-data processing (NDP) offer a promising path for expanding system memory to support DLRM training with large embedding tables. However, existing CXL+NDP designs are limited to single-GPU scenarios and fail to address key challenges in multi-GPU settings, like memory contention and embedding placement. To address this issue, we propose RecDM , an efficient training system for large-scale Rec ommendation models on D isaggregated M emory. RecDM features a modular, many-to-many CXL-based architecture with lightweight NDP units integrated into the CXL controller of each memory expansion unit, avoiding changes to DRAM chips or DIMM organization. To optimize training throughput, RecDM introduces a hierarchical memory device allocation strategy that balances memory bandwidth and capacity by combining shared and exclusive device mappings. Furthermore, RecDM proposes a bandwidth-driven 2D embedding table sharding method, enabling fine-grained placement across heterogeneous memory hierarchies and devices. Finally, RecDM incorporates an input-adaptive communication routing mechanism combined with a pipelined GPU-CXL execution model to reduce synchronization and data movement overhead. Comprehensive experiments demonstrate that RecDM achieves an average 10.2 × speedup over prior work that offloads embedding tables to host DRAM, and outperforms state-of-the-art CXL+NDP solution by an average of 2.2 ×.
Authors
- Yufei Ding (ORCID: https://orcid.org/0000-0002-8716-5793)
- Zhongkai Yu (ORCID: https://orcid.org/0000-0001-9893-3498)
- Yangwook Kang (ORCID: https://orcid.org/0009-0007-9536-1894)
- Yuke Wang (ORCID: https://orcid.org/0000-0001-5914-586X)
- Xulong Tang (ORCID: https://orcid.org/0000-0002-3385-2053)
- Liu Liu (ORCID: https://orcid.org/0000-0003-0792-8146)
- Yikai Li (ORCID: https://orcid.org/0000-0001-8943-2755)
- Zheng Wang (ORCID: https://orcid.org/0000-0002-8575-9432)
- Yichen Lin
- Kaijian Wang (ORCID: https://orcid.org/0009-0000-7036-8474)
Institutions
- Rensselaer Polytechnic Institute (US)
- University of Pittsburgh (US)
- University of California San Diego (US)
- Samsung Electronics (South Korea) (KR)
- Rice University (US)
Publication Details
- Journal
- ACM Transactions on Architecture and Code Optimization
- Published
- 2026-09-28
- DOI
- https://doi.org/10.1145/3848634
- Primary Topic
- Advanced Neural Network Applications
- Type
- article
- Field-Weighted Citation Impact
- 0.00