Contrastive Pretraining on Real LHC Jets: Transfer, Scorer Dependence and Mass Distortion

We compare contrastive pretraining on real AspenOpenJets and simulated LHC Olympics jets for background-based anomaly ranking, using matched random backbones and jet observables as controls. Equal training budgets comprise 100,000 jets per domain, one real-data shard and three seeds. Only 24 real-data training jets lie in the simulated reference sample’s central 90% transverse-momentum range, preventing a causal attribution to the pretraining domain. On partitions withheld from development, transverse-momentum balancing gives mean nearest-neighbor AUCs of 0.541, 0.564 and 0.517 for real-data, simulation and random representations on two-prong signals; three-prong values are 0.419, 0.553 and 0.426. All real-data seeds rank three-prong signal below background. Random backbones lead the original two-prong comparison with mean AUC 0.636. Signal–background momentum and multiplicity differences are consistent with score correlations and ranking changes, without establishing their cause. Regularized Mahalanobis and whitened nearest-neighbor scores do not make real-data representations competitive with girth; whitening benefits random features more. At background-calibrated operating points, girth retains more signal but strongly distorts background jet mass, whereas real-data-trained scores retain little signal. This bounded comparison characterizes representation, scorer and population dependence. Prior access to these public benchmarks limits independence; scorer and physics extensions are post-evaluation diagnosticsThis version supersedes the earlier preprint v1.0.0 (https://doi.org/10.5281/zenodo.20827792), which described a different analysis with different conclusions. Code and saved numerical results: https://github.com/Animesh-Parashar/aspen-jet-anomaly (archived at https://doi.org/10.5281/zenodo.23140427).

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-05
DOI
https://doi.org/10.5281/zenodo.20827791
Primary Topic
Particle physics theoretical and experimental studies
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

Contrastive Pretraining on Real LHC Jets: Transfer, Scorer Dependence and Mass Distortion

Animesh Parashar, Aditya Parashar
Zenodo (CERN European Organization for Nuclear Research)
Particle physics theoretical and experimental studies
preprint

Contrastive Pretraining on Real LHC Jets: Transfer, Scorer Dependence and Mass Distortion

Animesh Parashar, Aditya Parashar
preprint en

Abstract

We compare contrastive pretraining on real AspenOpenJets and simulated LHC Olympics jets for background-based anomaly ranking, using matched random backbones and jet observables as controls. Equal training budgets comprise 100,000 jets per domain, one real-data shard and three seeds. Only 24 real-data training jets lie in the simulated reference sample’s central 90% transverse-momentum range, preventing a causal attribution to the pretraining domain. On partitions withheld from development, transverse-momentum balancing gives mean nearest-neighbor AUCs of 0.541, 0.564 and 0.517 for real-data, simulation and random representations on two-prong signals; three-prong values are 0.419, 0.553 and 0.426. All real-data seeds rank three-prong signal below background. Random backbones lead the original two-prong comparison with mean AUC 0.636. Signal–background momentum and multiplicity differences are consistent with score correlations and ranking changes, without establishing their cause. Regularized Mahalanobis and whitened nearest-neighbor scores do not make real-data representations competitive with girth; whitening benefits random features more. At background-calibrated operating points, girth retains more signal but strongly distorts background jet mass, whereas real-data-trained scores retain little signal. This bounded comparison characterizes representation, scorer and population dependence. Prior access to these public benchmarks limits independence; scorer and physics extensions are post-evaluation diagnosticsThis version supersedes the earlier preprint v1.0.0 (https://doi.org/10.5281/zenodo.20827792), which described a different analysis with different conclusions. Code and saved numerical results: https://github.com/Animesh-Parashar/aspen-jet-anomaly (archived at https://doi.org/10.5281/zenodo.23140427).

Zenodo (CERN European Organization for Nuclear Research)
Birla Institute of Technology, Mesra (IN), Indian Institute of Technology Dhanbad (IN)
Particle physics theoretical and experimental studies
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.