EcoMLC: An Energy-Efficient Machine Learning Compilation Framework via Cache Reuse and Frequency Scaling
While modern machine learning compilers (MLCs) are pushing deep neural network (DNN) inference latency to its theoretical limits, the accompanying energy overhead has become a critical bottleneck for large-scale deployment. Excessive energy consumption not only increases electricity costs, but also affects the sustainability of edge platforms. Traditional performance-first optimizations often lead to significant energy waste, as they overlook the complex interplay between kernel configurations, cache behaviors and frequency scaling. This paper presents EcoMLC, an energy-efficient machine learning compilation framework. EcoMLC introduces a reverse scheduling policy combined with an explicit cache priority mechanism, mitigating cache thrashing inherent in standard least recently used (LRU) policies. EcoMLC employs an end-to-end frequency-sweeping analysis to identify optimal energy operating points and enables fine-grained trade-offs between performance and energy under diverse service level objective (SLO) constraints. Experimental results demonstrate that EcoMLC effectively optimizes various workloads, achieving an average of 43% and 26% energy savings (in J/request and J/token) over NVIDIA’s TensorRT and vLLM, and 22% energy savings on an AMD GPU compared to vLLM. For certain large language model (LLM) configurations, EcoMLC achieves a 33.9% reduction in energy consumption at the cost of a marginal 7% decline in performance.
Authors
- En Shao (ORCID: https://orcid.org/0000-0002-9678-7228)
- Leping Wang (ORCID: https://orcid.org/0009-0009-4940-5598)
- Dingwen Tao (ORCID: https://orcid.org/0000-0001-5422-4497)
- Zhanyuan Di (ORCID: https://orcid.org/0000-0003-2716-5051)
- Ninghui Sun (ORCID: https://orcid.org/0000-0002-1953-1392)
- Guangming Tan (ORCID: https://orcid.org/0000-0002-6361-5948)
- Chaoyang Gu (ORCID: https://orcid.org/0009-0007-5455-047X)
Institutions
- Institute of Computing Technology (CN)
- University of Chinese Academy of Sciences (CN)
Publication Details
- Journal
- ACM Transactions on Architecture and Code Optimization
- Published
- 2026-10-05
- DOI
- https://doi.org/10.1145/3856827
- Primary Topic
- Parallel Computing and Optimization Techniques
- Type
- article
- Field-Weighted Citation Impact
- 0.00