TCFG: A Two-Tier Collaborative Fine-Grained Programming Model Integrating On-Chip Multi-Asynchronous Unit Pipelining and Inter-Node Dynamic Scheduling
In the post-Moore’s Law era, many-core processors integrate diverse on-chip asynchronous accelerators. Existing programming models such as CUDA and OpenCL suffer from coarse-grained scheduling, and fine-grained optimizations only support limited computation-communication latency hiding. Moreover, intra-node asynchronous pipeline techniques are decoupled from cross-node dynamic scheduling runtimes, raising development costs for large-scale HPC applications. This paper proposes TCFG, a two-tier collaborative fine-grained programming model. Guided by compiler directives, it unifies on-chip accelerators as standardized pipelines with formal DAG-based stream abstraction and two adaptive buffer allocation strategies. It provides a layered declarative directive system for automatic loop transformation and dependency handling, and interfaces with existing dynamic scheduling runtimes through a unified runtime interface for joint intra-node pipelining and inter-node load balancing. Evaluations on a many-core simulator and SW26010Pro platform show up to 4.65× speedup for data-intensive workloads, 4.9–41.1% reduction in code size versus manual optimization, and 2.77× overall speedup for SVD tasks, of which 2.73× is achieved by inter-node dynamic scheduling alone and the remaining gain—a 1.5% reduction in execution time—by the intra-node asynchronous pipeline. Results demonstrate that TCFG delivers promising performance while lowering parallel-programming complexity.
Authors
- 董恩明
- Yiqing Liu (ORCID: https://orcid.org/0000-0003-2064-3938)
- Yanbing Li
- Qi Liu (ORCID: https://orcid.org/0009-0002-7141-4174)
- Yanfei Fang
- Yunfei Wang
- Xinhui Yuan
Publication Details
- Journal
- Electronics
- Published
- 2026-10-09
- DOI
- https://doi.org/10.3390/electronics15204592
- Primary Topic
- Parallel Computing and Optimization Techniques
- Type
- article
- Field-Weighted Citation Impact
- 0.00