Adaptive Thread Scheduling and Optimization Framework for the Multi-Core CPU with Large-Capacity Cache
The continuous growth in processor speed and the persistent memory latency bottleneck make large-capacity cache critical for mitigating the “memory wall”. However, topology-oblivious OS scheduling frequently induces severe cache contention and cross-die migrations, diminishing these hardware benefits. To address this, we propose a thread scheduling framework driven by runtime microarchitectural feedback. Instead of treating performance profiling and thread placement as isolated steps, our approach directly uses hardware performance counters to guide thread configurations to maximize cache utilization. First, the Thread and CPU Monitoring (TCM) framework tracks CPU and cache contention across both Simultaneous Multi-Threading (SMT) and non-SMT modes. Based on this feedback, the Thread Extension and Placement (TEP) algorithm dynamically adapts thread concurrency. To balance computational efficiency and runtime overhead, TEP determines the optimal thread count using cubic spline interpolation combined with a golden ratio search. Concurrently, a topology-aware affinity strategy binds threads to specific cores to maintain cache locality. Evaluations on AMD EPYC processors show that this approach improves overall execution efficiency. On average, the framework reduces the L3 miss rate from 62.71% to 51.89%, lowers CPI from 2.35 to 2.02, decreases cross-core thread migrations from 36902 to 33339, and reduces wall-clock time by 10.52%. The framework operates in user space without kernel modifications, ensuring ease of practical deployment for HPC and AI inference.
Authors
- Hongjun Dai (ORCID: https://orcid.org/0000-0002-1075-8750)
- Meikang Qiu (ORCID: https://orcid.org/0000-0002-1004-0140)
- Xinming Song
Institutions
- Shandong University (CN)
- Augusta University (US)
Publication Details
- Journal
- ACM Transactions on Architecture and Code Optimization
- Published
- 2026-10-09
- DOI
- https://doi.org/10.1145/3857768
- Primary Topic
- Parallel Computing and Optimization Techniques
- Type
- article
- Field-Weighted Citation Impact
- 0.00