TCFG: A Two-Tier Collaborative Fine-Grained Programming Model Integrating On-Chip Multi-Asynchronous Unit Pipelining and Inter-Node Dynamic Scheduling

In the post-Moore’s Law era, many-core processors integrate diverse on-chip asynchronous accelerators. Existing programming models such as CUDA and OpenCL suffer from coarse-grained scheduling, and fine-grained optimizations only support limited computation-communication latency hiding. Moreover, intra-node asynchronous pipeline techniques are decoupled from cross-node dynamic scheduling runtimes, raising development costs for large-scale HPC applications. This paper proposes TCFG, a two-tier collaborative fine-grained programming model. Guided by compiler directives, it unifies on-chip accelerators as standardized pipelines with formal DAG-based stream abstraction and two adaptive buffer allocation strategies. It provides a layered declarative directive system for automatic loop transformation and dependency handling, and interfaces with existing dynamic scheduling runtimes through a unified runtime interface for joint intra-node pipelining and inter-node load balancing. Evaluations on a many-core simulator and SW26010Pro platform show up to 4.65× speedup for data-intensive workloads, 4.9–41.1% reduction in code size versus manual optimization, and 2.77× overall speedup for SVD tasks, of which 2.73× is achieved by inter-node dynamic scheduling alone and the remaining gain—a 1.5% reduction in execution time—by the intra-node asynchronous pipeline. Results demonstrate that TCFG delivers promising performance while lowering parallel-programming complexity.

Authors

Publication Details

Journal
Electronics
Published
2026-10-09
DOI
https://doi.org/10.3390/electronics15204592
Primary Topic
Parallel Computing and Optimization Techniques
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

TCFG: A Two-Tier Collaborative Fine-Grained Programming Model Integrating On-Chip Multi-Asynchronous Unit Pipelining and Inter-Node Dynamic Scheduling

董恩明, Yiqing Liu, Yanbing Li, Qi Liu et al.
Electronics
Parallel Computing and Optimization Techniques
article

TCFG: A Two-Tier Collaborative Fine-Grained Programming Model Integrating On-Chip Multi-Asynchronous Unit Pipelining and Inter-Node Dynamic Scheduling

董恩明, Yiqing Liu, Yanbing Li, Qi Liu, Yanfei Fang, Yunfei Wang, Xinhui Yuan
article en

Abstract

In the post-Moore’s Law era, many-core processors integrate diverse on-chip asynchronous accelerators. Existing programming models such as CUDA and OpenCL suffer from coarse-grained scheduling, and fine-grained optimizations only support limited computation-communication latency hiding. Moreover, intra-node asynchronous pipeline techniques are decoupled from cross-node dynamic scheduling runtimes, raising development costs for large-scale HPC applications. This paper proposes TCFG, a two-tier collaborative fine-grained programming model. Guided by compiler directives, it unifies on-chip accelerators as standardized pipelines with formal DAG-based stream abstraction and two adaptive buffer allocation strategies. It provides a layered declarative directive system for automatic loop transformation and dependency handling, and interfaces with existing dynamic scheduling runtimes through a unified runtime interface for joint intra-node pipelining and inter-node load balancing. Evaluations on a many-core simulator and SW26010Pro platform show up to 4.65× speedup for data-intensive workloads, 4.9–41.1% reduction in code size versus manual optimization, and 2.77× overall speedup for SVD tasks, of which 2.73× is achieved by inter-node dynamic scheduling alone and the remaining gain—a 1.5% reduction in execution time—by the intra-node asynchronous pipeline. Results demonstrate that TCFG delivers promising performance while lowering parallel-programming complexity.

ElectronicsVol. 15(20)
Openalex Percentile: Top 8%
Parallel Computing and Optimization Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

TCFG: A Two-Tier Collaborative Fine-Grained Programming Model Integrating On-Chip Multi-Asynchronous Unit Pipelining and Inter-Node Dynamic Scheduling — 董恩明, Yiqing Liu, et al. · Electronics (2026) | TGRS Research Map | TGRS