Zheng Wang, Zhongkai Yu, Kaijian Wang, Yichen Lin, Yikai Li, Liu Liu, Xulong Tang, Yuke Wang, Yangwook Kang, Yufei Ding · ACM Transactions on Architecture and Code Optimization 2026 · 2026
DOI: 10.1145/3848634
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
The embedding tables in Deep Learning Recommendation Models (DLRMs) require significant memory capacity and bandwidth but relatively lower computing power, making it economically inefficient to scale by adding more GPUs solely to meet memory requirements. Recent advances in Compute Express Link (CXL) and near-data processing (NDP) offer a promising path for expanding system memory to support DLRM training with large embedding tables. However, existing CXL+NDP designs are limited to single-GPU scenarios and fail to address key challenges in multi-GPU settings, like memory contention and embedding placement. To address this issue, we propose RecDM , an efficient training system for large-scale Rec ommendation models on D isaggregated M emory. RecDM features a modular, many-to-many CXL-based architecture with lightweight NDP units integrated into the CXL controller of each memory expansion unit, avoiding changes to DRAM chips or DIMM organization. To optimize training throughput, RecDM introduces a hierarchical memory device allocation strategy that balances memory bandwidth and capacity by combining shared and exclusive device mappings. Furthermore, RecDM proposes a bandwidth-driven 2D embedding table sharding method, enabling fine-grained placement across heterogeneous memory hierarchies and devices. Finally, RecDM incorporates an input-adaptive communication routing mechanism combined with a pipelined GPU-CXL execution model to reduce synchronization and data movement overhead. Comprehensive experiments demonstrate that RecDM achieves an average 10.2 × speedup over prior work that offloads embedding tables to host DRAM, and outperforms state-of-the-art CXL+NDP solution by an average of 2.2 ×.
No comments yet — start the discussion below.