Adaptive and Fault-Tolerant Scheduling Framework for Scalable and High-Performance Distributed Systems with Dynamic and Heterogeneous Workloads

Abstract High-performance distributed systems (HPDS) are increasingly challenged by dynamic workloads, heterogeneous resources, and susceptibility to faults, which degrade scheduling efficiency and resilience. Traditional heuristic and learning-based scheduling approaches, on the other hand, tend to be poorly adaptable and have high overhead in large-scale deployments. To solve these issues, this paper proposes an Adaptive Quantum Neuro-Symbolic Scheduling and Federated Sparse Transformer (AQNSS-FST) framework that is augmented with Harris Hawks Optimization (HHO). The AQNSS component employs quantum-inspired symbolic reasoning for adaptive task-to-resource mapping, in which quantum operators boost convergence in high-dimensional scheduling spaces as compared to classical heuristics. The FST module facilitates collaborative, lightweight, and fault-resilient learning across distributed nodes, and HHO dynamically balances global exploration and local exploitation to avoid premature convergence under varying workloads. The experimental results on the Google Cluster Workload Traces and FEMNIST dataset show that AQNSS-FST achieves a significantly improved performance compared to the current baselines. The structure saves up to 25% of makespan, 20–25% of throughput, and a nearly 10–15% improvement in reliability of task completion with Fault Tolerance Rate (FTR) and Task Completion Rate (TCR) gains. Moreover, convergence rounds are reduced by over 30%, and runtime overhead is minimized, as well as the average iteration time is 1.8 s as opposed to 3.8 s of heuristics. These results lead to the conclusion that a scalable and fault-tolerant solution for next-generation HPDS can be achieved by integrating quantum-inspired neuro-symbolic scheduling, federated sparse learning, and HHO-driven optimization.

Authors

Institutions

Publication Details

Journal
Mobile Networks and Applications
Published
2026-10-09
DOI
https://doi.org/10.1007/s11036-026-02530-8
Primary Topic
Distributed and Parallel Computing Systems
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Adaptive and Fault-Tolerant Scheduling Framework for Scalable and High-Performance Distributed Systems with Dynamic and Heterogeneous Workloads

Khalid K. Almuzaini, Piyush Kumar Shukla, Muhammad Faheem, Kailin Yang
Mobile Networks and Applications
Distributed and Parallel Computing Systems
article

Adaptive and Fault-Tolerant Scheduling Framework for Scalable and High-Performance Distributed Systems with Dynamic and Heterogeneous Workloads

Khalid K. Almuzaini, Piyush Kumar Shukla, Muhammad Faheem, Kailin Yang
article en

Abstract

Abstract High-performance distributed systems (HPDS) are increasingly challenged by dynamic workloads, heterogeneous resources, and susceptibility to faults, which degrade scheduling efficiency and resilience. Traditional heuristic and learning-based scheduling approaches, on the other hand, tend to be poorly adaptable and have high overhead in large-scale deployments. To solve these issues, this paper proposes an Adaptive Quantum Neuro-Symbolic Scheduling and Federated Sparse Transformer (AQNSS-FST) framework that is augmented with Harris Hawks Optimization (HHO). The AQNSS component employs quantum-inspired symbolic reasoning for adaptive task-to-resource mapping, in which quantum operators boost convergence in high-dimensional scheduling spaces as compared to classical heuristics. The FST module facilitates collaborative, lightweight, and fault-resilient learning across distributed nodes, and HHO dynamically balances global exploration and local exploitation to avoid premature convergence under varying workloads. The experimental results on the Google Cluster Workload Traces and FEMNIST dataset show that AQNSS-FST achieves a significantly improved performance compared to the current baselines. The structure saves up to 25% of makespan, 20–25% of throughput, and a nearly 10–15% improvement in reliability of task completion with Fault Tolerance Rate (FTR) and Task Completion Rate (TCR) gains. Moreover, convergence rounds are reduced by over 30%, and runtime overhead is minimized, as well as the average iteration time is 1.8 s as opposed to 3.8 s of heuristics. These results lead to the conclusion that a scalable and fault-tolerant solution for next-generation HPDS can be achieved by integrating quantum-inspired neuro-symbolic scheduling, federated sparse learning, and HHO-driven optimization.

Mobile Networks and Applications
King Abdulaziz City for Science and Technology (SA), Institute of Management Technology (IN), Rajiv Gandhi Technical University (IN), Sichuan Technology and Business University (CN), VTT Technical Research Centre of Finland (FI)
Openalex Percentile: Top 11%
Distributed and Parallel Computing Systems
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.