Task aggregation as a strategy to optimize Earth System Model workflows in HPC: assessing real scenarios with EC-Earth

Earth System Models (ESMs) are commonly executed as complex workflows consisting of numerous interdependent tasks that comprise steps such as model and data deployment, simulation, data transfer, and post-processing. Workflows facilitate the execution of long high-resolution configurations by splitting the runtime of the simulation into different tasks, in order to ensure frequent checkpointing and to comply with the restrictions of the large High-Performance Computing (HPC) machines where they are executed. These machines are frequently congested due to their high demand and, therefore, implement scheduling policies to share the resources among the users. This complicates the execution of long ensemble simulation workflows, where each successive simulation task can only be submitted once the preceding one finishes, causing queue time to accumulate, which is the duration for which the jobs wait for the HPC platform to allocate the required resources for their execution. To alleviate this issue, we propose achieving shorter times-to-response, which are the intervals from the first submission to the completion of the final task, by applying a solution to reduce subsequent requests for resources and, consequently, reducing queue times. This solution, known as task aggregation, consists of grouping multiple tasks and submitting them as a single job, respecting their dependencies, and without altering their underlying logic. This technique separates the workflow definition from how individual tasks are submitted and executed, allowing the workflow manager to make decisions during runtime. In this paper, we performed the first controlled assessment on the effects of task aggregation by conducting concurrent pairs of climate simulations in production machines, with the sole difference that one uses aggregation. These simulations were executed on three European supercomputers: MeluXina, MareNostrum 4, and MareNostrum 5. We measured the evolution of the fair share, a scheduling factor that normally plays a major role in the priority of the jobs. Besides absolute gains, we compute the impact of aggregation by comparing consolidated performance metrics for climate models within the community. We prove the benefits of the task aggregation using EC-Earth3, a widely used European community climate model that shares main features with many other ESMs and has a representative workload. The experimental findings of our research indicate that across the three evaluated supercomputing platforms, applying task aggregation decreases total queue times by 11.17 to 12.33 times compared to a workflow that does not , representing an improvement of up to 23 % more years simulated in the same time span in the case of the platform with the highest congestion. Therefore, results have shown that task aggregation proves to be beneficial for long climate simulations. Moreover, we have credible reasons to believe that any vertical (also called chained) workflow should benefit from using it. We explain that this reduction in the time-to-solution comes from the decrease in the number of submitted jobs and the utilization of the machine. By aggregating tasks, we had many times less jobs queued and, albeit their longer length, we observed that they are in queue less time than if they had been submitted individually. Our results also show that the user's jobs are being held in queue in spite of their utilization due to the fair share being influenced by the other members of the HPC account, which is a direct consequence of the fair share policy of the machine.

Authors

Institutions

Publication Details

Journal
Geoscientific model development
Published
2026-09-11
DOI
https://doi.org/10.5194/gmd-19-8501-2026
Primary Topic
Distributed and Parallel Computing Systems
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Task aggregation as a strategy to optimize Earth System Model workflows in HPC: assessing real scenarios with EC-Earth

Miguel Castrillo, Mario Acosta, Manuel Giménez de Castro Marciani, Pablo Goitia González
Geoscientific model development
Distributed and Parallel Computing Systems
article

Task aggregation as a strategy to optimize Earth System Model workflows in HPC: assessing real scenarios with EC-Earth

Miguel Castrillo, Mario Acosta, Manuel Giménez de Castro Marciani, Pablo Goitia González
article en

Abstract

Earth System Models (ESMs) are commonly executed as complex workflows consisting of numerous interdependent tasks that comprise steps such as model and data deployment, simulation, data transfer, and post-processing. Workflows facilitate the execution of long high-resolution configurations by splitting the runtime of the simulation into different tasks, in order to ensure frequent checkpointing and to comply with the restrictions of the large High-Performance Computing (HPC) machines where they are executed. These machines are frequently congested due to their high demand and, therefore, implement scheduling policies to share the resources among the users. This complicates the execution of long ensemble simulation workflows, where each successive simulation task can only be submitted once the preceding one finishes, causing queue time to accumulate, which is the duration for which the jobs wait for the HPC platform to allocate the required resources for their execution. To alleviate this issue, we propose achieving shorter times-to-response, which are the intervals from the first submission to the completion of the final task, by applying a solution to reduce subsequent requests for resources and, consequently, reducing queue times. This solution, known as task aggregation, consists of grouping multiple tasks and submitting them as a single job, respecting their dependencies, and without altering their underlying logic. This technique separates the workflow definition from how individual tasks are submitted and executed, allowing the workflow manager to make decisions during runtime. In this paper, we performed the first controlled assessment on the effects of task aggregation by conducting concurrent pairs of climate simulations in production machines, with the sole difference that one uses aggregation. These simulations were executed on three European supercomputers: MeluXina, MareNostrum 4, and MareNostrum 5. We measured the evolution of the fair share, a scheduling factor that normally plays a major role in the priority of the jobs. Besides absolute gains, we compute the impact of aggregation by comparing consolidated performance metrics for climate models within the community. We prove the benefits of the task aggregation using EC-Earth3, a widely used European community climate model that shares main features with many other ESMs and has a representative workload. The experimental findings of our research indicate that across the three evaluated supercomputing platforms, applying task aggregation decreases total queue times by 11.17 to 12.33 times compared to a workflow that does not , representing an improvement of up to 23 % more years simulated in the same time span in the case of the platform with the highest congestion. Therefore, results have shown that task aggregation proves to be beneficial for long climate simulations. Moreover, we have credible reasons to believe that any vertical (also called chained) workflow should benefit from using it. We explain that this reduction in the time-to-solution comes from the decrease in the number of submitted jobs and the utilization of the machine. By aggregating tasks, we had many times less jobs queued and, albeit their longer length, we observed that they are in queue less time than if they had been submitted individually. Our results also show that the user's jobs are being held in queue in spite of their utilization due to the fair share being influenced by the other members of the HPC account, which is a direct consequence of the fair share policy of the machine.

Geoscientific model developmentVol. 19(17)
Universidad de Cantabria (ES), Barcelona Supercomputing Center (ES), Universitat Politècnica de Catalunya (ES)
Consejo Superior de Investigaciones Científicas
Climate action
Openalex Percentile: Top 8%
Distributed and Parallel Computing Systems
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.