Empirical Scaling Laws of Quantum Vision Transformers
Quantum machine learning (QML) is a promising application of quantum computing, yet when and how practical learning models can achieve quantum advantage remains poorly understood. Even a polynomial improvement in the resources needed to attain a given predictive performance would be significant. In this work, we investigate empirical scaling laws of quantum vision transformers (QViTs), which connect QML to the transformer architecture underlying modern large language models. We ask whether QViTs exhibit scaling behavior analogous to classical models and how this behavior changes when qubit number is included alongside dataset size and model parameters. The validation losses in our numerical image-classification experiments are well described by a joint shifted power-law model in qubit number, transformer depth, and training dataset size. Expressing this fit in terms of the classical encoder parameter budget identifies a regime in which increasing the number of qubits can compensate for fewer classical encoder parameters at comparable predicted performance. This finding is consistent with earlier observations of parameter efficiency in quantum learning models and suggests a direction for exploring possible resource savings in QML.
Authors
- Fengyi Gao (ORCID: https://orcid.org/0009-0004-9169-2192)
- Junyu Liu (ORCID: https://orcid.org/0000-0003-1669-8039)
- Qi Cheng (ORCID: https://orcid.org/0009-0000-8764-7349)
- Xin Jin
- Xiaowei Jia (ORCID: https://orcid.org/0000-0001-8544-5233)
- Yuqing Li
Institutions
- Rutgers, The State University of New Jersey (US)
- University of Pittsburgh (US)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-28
- DOI
- https://doi.org/10.5281/zenodo.23005853
- Primary Topic
- Quantum Computing Algorithms and Architecture
- Type
- preprint