Feature Engineering for Queue Waiting Time Prediction: Temporal and Queue-State Reconstruction on the Theta Supercomputer
Predicting queue waiting time for batch job schedulers in high-performance computing (HPC) systems is a critical research topic aimed at maximizing resource utilization efficiency and enhancing user experience. However, existing prediction approaches heavily rely on static job characteristics provided by users at submission time, exhibiting clear limitations in capturing the dynamic congestion state of the scheduling system. To address these limitations, a log-based feature engineering approach designed to significantly improve queue waiting time prediction accuracy is proposed in this study. Specifically, feature redundancy is eliminated through comparative analysis, and temporal features capturing periodic job submission patterns are generated. Furthermore, global and local (queue-level) waiting and running states at job submission time are reconstructed exclusively from historical job logs, thereby representing dynamic queue congestion in a fine-grained manner. The effectiveness of the proposed framework was validated using a large-scale job log dataset collected from the Theta supercomputer at the Argonne Leadership Computing Facility (ALCF) between 2017 and 2023, where queue waiting time was discretized into nine distinct intervals and formulated as a multi-class classification problem. Finally, an ablation study was conducted to quantitatively evaluate the performance contribution of each engineered feature group, and comparative experiments were carried out across multiple machine learning models, including Naïve Bayes, k-nearest neighbors (KNN), Random Forest, and XGBoost. Among the compared models, Random Forest achieved the best overall performance, with an accuracy of 88.89%, a balanced accuracy of 57.65%, a macro-averaged precision of 67.95%, and a macro-averaged F1-score of 61.71%. Compared with the baseline feature set, the proposed feature engineering strategy improved accuracy, balanced accuracy, and macro F1-score by 3.85, 7.28, and 9.82 percentage points, respectively, and the proposed RF model obtained a balanced accuracy 18 percentage points higher than the best-performing model reported in a comparable reference study on the same dataset, a cross-study difference not directly attributable to the proposed feature engineering alone.
Authors
- Ju-Won Park (ORCID: https://orcid.org/0000-0003-1388-1583)
- Tsatsral Amarbayasgalan (ORCID: https://orcid.org/0000-0001-8399-655X)
- Eunhye Kim
Institutions
- Kunsan National University (KR)
- Electronics and Telecommunications Research Institute (KR)
- Korea Atomic Energy Research Institute (KR)
Publication Details
- Journal
- Applied Sciences
- Published
- 2026-09-09
- DOI
- https://doi.org/10.3390/app16188965
- Primary Topic
- Cloud Computing and Resource Management
- Type
- article
- Field-Weighted Citation Impact
- 0.00