Adaptive Thread Scheduling and Optimization Framework for the Multi-Core CPU with Large-Capacity Cache

The continuous growth in processor speed and the persistent memory latency bottleneck make large-capacity cache critical for mitigating the “memory wall”. However, topology-oblivious OS scheduling frequently induces severe cache contention and cross-die migrations, diminishing these hardware benefits. To address this, we propose a thread scheduling framework driven by runtime microarchitectural feedback. Instead of treating performance profiling and thread placement as isolated steps, our approach directly uses hardware performance counters to guide thread configurations to maximize cache utilization. First, the Thread and CPU Monitoring (TCM) framework tracks CPU and cache contention across both Simultaneous Multi-Threading (SMT) and non-SMT modes. Based on this feedback, the Thread Extension and Placement (TEP) algorithm dynamically adapts thread concurrency. To balance computational efficiency and runtime overhead, TEP determines the optimal thread count using cubic spline interpolation combined with a golden ratio search. Concurrently, a topology-aware affinity strategy binds threads to specific cores to maintain cache locality. Evaluations on AMD EPYC processors show that this approach improves overall execution efficiency. On average, the framework reduces the L3 miss rate from 62.71% to 51.89%, lowers CPI from 2.35 to 2.02, decreases cross-core thread migrations from 36902 to 33339, and reduces wall-clock time by 10.52%. The framework operates in user space without kernel modifications, ensuring ease of practical deployment for HPC and AI inference.

Authors

Institutions

Publication Details

Journal
ACM Transactions on Architecture and Code Optimization
Published
2026-10-09
DOI
https://doi.org/10.1145/3857768
Primary Topic
Parallel Computing and Optimization Techniques
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Adaptive Thread Scheduling and Optimization Framework for the Multi-Core CPU with Large-Capacity Cache

Hongjun Dai, Meikang Qiu, Xinming Song
ACM Transactions on Architecture and Code Optimization
Parallel Computing and Optimization Techniques
article

Adaptive Thread Scheduling and Optimization Framework for the Multi-Core CPU with Large-Capacity Cache

Hongjun Dai, Meikang Qiu, Xinming Song
article en

Abstract

The continuous growth in processor speed and the persistent memory latency bottleneck make large-capacity cache critical for mitigating the “memory wall”. However, topology-oblivious OS scheduling frequently induces severe cache contention and cross-die migrations, diminishing these hardware benefits. To address this, we propose a thread scheduling framework driven by runtime microarchitectural feedback. Instead of treating performance profiling and thread placement as isolated steps, our approach directly uses hardware performance counters to guide thread configurations to maximize cache utilization. First, the Thread and CPU Monitoring (TCM) framework tracks CPU and cache contention across both Simultaneous Multi-Threading (SMT) and non-SMT modes. Based on this feedback, the Thread Extension and Placement (TEP) algorithm dynamically adapts thread concurrency. To balance computational efficiency and runtime overhead, TEP determines the optimal thread count using cubic spline interpolation combined with a golden ratio search. Concurrently, a topology-aware affinity strategy binds threads to specific cores to maintain cache locality. Evaluations on AMD EPYC processors show that this approach improves overall execution efficiency. On average, the framework reduces the L3 miss rate from 62.71% to 51.89%, lowers CPI from 2.35 to 2.02, decreases cross-core thread migrations from 36902 to 33339, and reduces wall-clock time by 10.52%. The framework operates in user space without kernel modifications, ensuring ease of practical deployment for HPC and AI inference.

ACM Transactions on Architecture and Code Optimization
Shandong University (CN), Augusta University (US)
Openalex Percentile: Top 8%
Parallel Computing and Optimization Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Adaptive Thread Scheduling and Optimization Framework for the Multi-Core CPU with Large-Capacity Cache — Hongjun Dai, Meikang Qiu, et al. · ACM Transactions on Architecture and Code Optimization (2026) | TGRS Research Map | TGRS