FCoral: A Framework for Fine-grained Concurrency Evaluation Across Programming Language Runtimes
Modern programming languages increasingly support lightweight concurrency abstractions (e.g., coroutines), which are managed by language-level runtimes to improve scalability and performance in highly concurrent applications. Compared with coarse-grained thread-level concurrency, the performance of fine-grained coroutine-level concurrency depends not only on system resources but is even more sensitive to the efficiency of the concurrent runtimes. Performance evaluation across concurrent runtimes is important for understanding performance bottlenecks and improving the efficiency of both concurrent applications and runtime systems. However, existing approaches mainly rely on manually crafted, language-specific benchmarks, making it difficult to ensure semantic consistency across languages and perform fair cross-runtime comparisons. Moreover, they often treat concurrent runtimes as black boxes, overlooking the impact of internal runtime mechanisms on concurrency performance. In this paper, we propose FCoral, a cross-language framework for fine-grained concurrency evaluation. By constructing formally verified and language-agnostic Concurrent Workload Programs (CWPs), FCoral automatically generates semantically consistent benchmarks across different runtimes. It further employs a parameterized methodology to evaluate runtime performance across different concurrency scales, task granularities, and workload types. Using FCoral, we conduct a comprehensive study of five representative runtimes–—FFRT (C++), Go, JVM, Tokio (Rust), and Cangjie–—across CPU-intensive, I/O-intensive, and mixed workloads. We further evaluate FCoral on two real-world high-concurrency applications. Results show that FCoral successfully constructs cross-language benchmarks, facilitating systematic performance comparisons across concurrent runtimes. Based on the evaluation results, we propose two application-level optimization strategies. Experimental results demonstrate that these strategies not only achieve significant performance improvements over the unoptimized versions but also further validate the effectiveness of FCoral.
Authors
- Weixing Ji (ORCID: https://orcid.org/0000-0002-3250-0435)
- Jianhua Gao (ORCID: https://orcid.org/0000-0002-3828-0015)
- Yuxiang Zhang (ORCID: https://orcid.org/0000-0003-2044-1199)
- Danying Ge (ORCID: https://orcid.org/0009-0007-4363-7625)
- Jianjun Shi (ORCID: https://orcid.org/0009-0007-5641-6354)
- Bingxin Liu (ORCID: https://orcid.org/0009-0005-4361-009X)
- Yinghui Huang (ORCID: https://orcid.org/0009-0004-1430-7211)
Institutions
- Beijing Normal University (CN)
Publication Details
- Journal
- ACM Transactions on Architecture and Code Optimization
- Published
- 2026-10-03
- DOI
- https://doi.org/10.1145/3847668
- Primary Topic
- Parallel Computing and Optimization Techniques
- Type
- article
- Field-Weighted Citation Impact
- 0.00