Adaptive and Fault-Tolerant Scheduling Framework for Scalable and High-Performance Distributed Systems with Dynamic and Heterogeneous Workloads
Abstract High-performance distributed systems (HPDS) are increasingly challenged by dynamic workloads, heterogeneous resources, and susceptibility to faults, which degrade scheduling efficiency and resilience. Traditional heuristic and learning-based scheduling approaches, on the other hand, tend to be poorly adaptable and have high overhead in large-scale deployments. To solve these issues, this paper proposes an Adaptive Quantum Neuro-Symbolic Scheduling and Federated Sparse Transformer (AQNSS-FST) framework that is augmented with Harris Hawks Optimization (HHO). The AQNSS component employs quantum-inspired symbolic reasoning for adaptive task-to-resource mapping, in which quantum operators boost convergence in high-dimensional scheduling spaces as compared to classical heuristics. The FST module facilitates collaborative, lightweight, and fault-resilient learning across distributed nodes, and HHO dynamically balances global exploration and local exploitation to avoid premature convergence under varying workloads. The experimental results on the Google Cluster Workload Traces and FEMNIST dataset show that AQNSS-FST achieves a significantly improved performance compared to the current baselines. The structure saves up to 25% of makespan, 20–25% of throughput, and a nearly 10–15% improvement in reliability of task completion with Fault Tolerance Rate (FTR) and Task Completion Rate (TCR) gains. Moreover, convergence rounds are reduced by over 30%, and runtime overhead is minimized, as well as the average iteration time is 1.8 s as opposed to 3.8 s of heuristics. These results lead to the conclusion that a scalable and fault-tolerant solution for next-generation HPDS can be achieved by integrating quantum-inspired neuro-symbolic scheduling, federated sparse learning, and HHO-driven optimization.
Authors
- Khalid K. Almuzaini (ORCID: https://orcid.org/0000-0003-4458-3730)
- Piyush Kumar Shukla (ORCID: https://orcid.org/0000-0002-3715-3882)
- Muhammad Faheem
- Kailin Yang
Institutions
- King Abdulaziz City for Science and Technology (SA)
- Institute of Management Technology (IN)
- Rajiv Gandhi Technical University (IN)
- Sichuan Technology and Business University (CN)
- VTT Technical Research Centre of Finland (FI)
Publication Details
- Journal
- Mobile Networks and Applications
- Published
- 2026-10-09
- DOI
- https://doi.org/10.1007/s11036-026-02530-8
- Primary Topic
- Distributed and Parallel Computing Systems
- Type
- article
- Field-Weighted Citation Impact
- 0.00