DPSM: Fine-Grained GPU Resource Management via Kernel Interception and TPC Control
Modern deep learning workloads are increasingly co-located on a single GPU to improve hardware utilization, yet existing MPS-based sharing mechanisms expose concurrency without providing QoS-aware control. Deadline-constrained training jobs can suffer severe slowdown when co-executed with best-effort workloads, because MPS does not reason about job urgency or deadline slack. This paper presents DPSM, a transparent runtime system for QoS-first, throughput-aware co-scheduling of concurrent deep learning jobs on a single NVIDIA GPU. DPSM does not replace the native GPU scheduler, generate deadlines, or assign deadlines to individual kernels. Instead, it intercepts CUDA kernel launches, tracks job-level progress and slowdown online, and identifies the most urgent job using a slack-based scoring model. DPSM then enforces the scheduling decision through two complementary mechanisms: launch-timing control for large-grid best-effort kernels, and TPC-mask control that limits the TPC budget of selected small-grid best-effort kernels. Under MPS, DPSM treats TPC masks as budget controls that limit the number of available TPCs. This design preserves the MPS execution model while mitigating temporal and resource-budget interference among co-running jobs. We evaluate DPSM on NVIDIA RTX A6000 GPUs using representative two-job and three-job deep learning co-location workloads. Across 1260 job-deadline instances, DPSM achieves the highest all-job deadline completion rate (DCR) of 39.2%, compared with 37.6% for TGS, 28.7% for MPS, and 18.6% for Orion. DPSM also maintains competitive normalized iteration throughput of 0.471, close to the highest-throughput baseline Orion at 0.499, while providing substantially better deadline satisfaction. We envision DPSM as a valuable tool for the community and have open-sourced it to facilitate future research at https://github.com/HIT-CeeCG/DPSM.
Authors
- Sichao Chen (ORCID: https://orcid.org/0009-0002-5696-4944)
- Hongwei Yang (ORCID: https://orcid.org/0000-0002-8386-0131)
- Meng Hao (ORCID: https://orcid.org/0000-0003-0043-4370)
- Shuo Si (ORCID: https://orcid.org/0009-0006-7028-5779)
- Wei Zhang (ORCID: https://orcid.org/0000-0001-7800-3189)
- Desheng Wang (ORCID: https://orcid.org/0000-0002-7502-7094)
- Fuzhen Yong (ORCID: https://orcid.org/0009-0009-2737-724X)
Institutions
- Harbin Institute of Technology (CN)
Publication Details
- Journal
- ACM Transactions on Architecture and Code Optimization
- Published
- 2026-10-08
- DOI
- https://doi.org/10.1145/3857801
- Primary Topic
- Parallel Computing and Optimization Techniques
- Type
- article
- Field-Weighted Citation Impact
- 0.00