Parallel and Distributed Training of Variational Quantum Algorithms on NISQ Systems: A Systematic Review of Scaling, Noise, and Cost
Background: Parallel and distributed execution is proposed to relieve throughput, trainability and resource constraints in variational quantum algorithms (VQAs), but the label “parallel” covers physically and statistically different interventions. We assessed what has actually been measured on noisy intermediate-scale quantum (NISQ) systems and whether reported benefits survive a consistent accounting of topology, baseline, noise and cost. Methods: We conducted a PRISMA-aligned systematic review using reproducible Scopus and Web of Science searches, supplementary semantic discovery in Elicit, a Zotero catalogue and supplied full texts. Elicit was searched on 17 February 2026 as a supplementary semantic-discovery source. The search returned 50 results. Ten records were bibliographically identifiable in the retained Elicit report. Because the original complete export is unavailable, the remaining 40 results cannot be individually identified, deduplicated, screened, or assigned study-level exclusion reasons. Elicit was therefore treated as a supplementary discovery source rather than the sole reproducible database-search mechanism. Eligible evidence was classified as Core, Supporting or Background. Seventeen Core studies were synthesized by mechanism and research question; no meta-analysis was performed because execution topologies, denominators and outcomes were incompatible. Results: The sources produced 207 appearances: 50 Elicit hits, 10 Scopus records, 29 Web of Science records, 47 Zotero records and 71 supplied full-text reports. Removing the 40 unidentifiable Elicit hits left 167 named and traceable appearances; 40 known duplicate or superseded appearances were removed, leaving 127 unique identifiable records. Fifty-seven were excluded and 70 were retained as 17 Core, 45 Supporting and 8 Background records. Final classification agreement was 127/127. Reported benefits comprised measured multi-device ensemble completion, single-QPU spatial throughput, simulated or modeled acceleration, measurement or iteration reduction and quality improvement. The most informative hardware evidence included a 10.5× mean multi-device ensemble speedup, a 35–45× single-QPU forward-step acceleration, an 18× single-QPU heatmap timing result, a derived 6× BayesMGD result and an SPSA slowdown to 0.58×. These values were not pooled. Queue, communication, retry, monetary and energy boundaries were usually incomplete. Conclusions: Parallel and distributed VQA training is technically plural rather than a single scalable intervention. Measured benefit depended on the work partition, device topology, quality target and cost boundary; favorable factors from different denominators cannot establish general speedup. The review contributes an unvalidated minimum reporting standard, a cost-boundary accounting framework and a conceptual cost-aware orchestration architecture. Matched end-to-end multi-QPU experiments, negative-result reporting and reproducible cost logs are required before performance claims can be generalized. Keywords: variational quantum algorithm; distributed quantum computing; parallel quantum computing; quantum machine learning; NISQ; quantum federated learning; systematic review; cost accounting
Authors
- Wael Al-Thohli
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-15
- DOI
- https://doi.org/10.5281/zenodo.22764424
- Primary Topic
- Quantum Computing Algorithms and Architecture
- Type
- article
- Field-Weighted Citation Impact
- 0.00