Scaling behaviour of transformer-based jet flavour tagging with the CMS experiment at ¿s = 13.6 TeV

Neural scaling laws describe empirical relationships between machine-learning performance and the resources used for training, including model capacity, dataset size, and computational cost. Such behaviour has been extensively studied for large language models and has recently been observed in jet-classification tasks. This note presents a systematic study of scaling behaviour for transformer-based jet flavour tagging using CMS simulation at ¿s=13.6 TeV. A family of Particle Transformer models is trained while varying model capacity, training-dataset size, and architectural configurations. The training dataset contains up to 301 million jets after flavour and kinematic reweighting. Training, validation, and test losses decrease systematically with increasing scale, and the corresponding gains translate into improved b- and c-jet tagging performance. Quark/gluon discrimination exhibits substantially weaker scaling. The interplay between model size, dataset size, and training compute is studied to identify efficient scaling regimes. The results demonstrate substantial remaining performance potential from increasing the scale of jet-tagging models, while quantifying the rapidly increasing computational cost required to obtain these improvements.

Authors

Publication Details

Journal
CERN Document Server (European Organization for Nuclear Research)
Published
2026-09-14
Primary Topic
Particle physics theoretical and experimental studies
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Scaling behaviour of transformer-based jet flavour tagging with the CMS experiment at ¿s = 13.6 TeV

CMS Collaboration
CERN Document Server (European Organization for Nuclear Research)
Particle physics theoretical and experimental studies
article

Scaling behaviour of transformer-based jet flavour tagging with the CMS experiment at ¿s = 13.6 TeV

CMS Collaboration
article en

Abstract

Neural scaling laws describe empirical relationships between machine-learning performance and the resources used for training, including model capacity, dataset size, and computational cost. Such behaviour has been extensively studied for large language models and has recently been observed in jet-classification tasks. This note presents a systematic study of scaling behaviour for transformer-based jet flavour tagging using CMS simulation at ¿s=13.6 TeV. A family of Particle Transformer models is trained while varying model capacity, training-dataset size, and architectural configurations. The training dataset contains up to 301 million jets after flavour and kinematic reweighting. Training, validation, and test losses decrease systematically with increasing scale, and the corresponding gains translate into improved b- and c-jet tagging performance. Quark/gluon discrimination exhibits substantially weaker scaling. The interplay between model size, dataset size, and training compute is studied to identify efficient scaling regimes. The results demonstrate substantial remaining performance potential from increasing the scale of jet-tagging models, while quantifying the rapidly increasing computational cost required to obtain these improvements.

CERN Document Server (European Organization for Nuclear Research)
Peace, Justice and strong institutions
Openalex Percentile: Top 12%
Particle physics theoretical and experimental studies
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Scaling behaviour of transformer-based jet flavour tagging with the CMS experiment at ¿s = 13.6 TeV — CMS Collaboration · CERN Document Server (European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS