Scaling behaviour of transformer-based jet flavour tagging with the CMS experiment at ¿s = 13.6 TeV
Neural scaling laws describe empirical relationships between machine-learning performance and the resources used for training, including model capacity, dataset size, and computational cost. Such behaviour has been extensively studied for large language models and has recently been observed in jet-classification tasks. This note presents a systematic study of scaling behaviour for transformer-based jet flavour tagging using CMS simulation at ¿s=13.6 TeV. A family of Particle Transformer models is trained while varying model capacity, training-dataset size, and architectural configurations. The training dataset contains up to 301 million jets after flavour and kinematic reweighting. Training, validation, and test losses decrease systematically with increasing scale, and the corresponding gains translate into improved b- and c-jet tagging performance. Quark/gluon discrimination exhibits substantially weaker scaling. The interplay between model size, dataset size, and training compute is studied to identify efficient scaling regimes. The results demonstrate substantial remaining performance potential from increasing the scale of jet-tagging models, while quantifying the rapidly increasing computational cost required to obtain these improvements.
Authors
- CMS Collaboration
Publication Details
- Journal
- CERN Document Server (European Organization for Nuclear Research)
- Published
- 2026-09-14
- Primary Topic
- Particle physics theoretical and experimental studies
- Type
- article
- Field-Weighted Citation Impact
- 0.00