Speaker
Description
Foundation models for collider physics have so far been trained predominantly on simulated events or proton–proton collision data, leaving their behaviour in the high-occupancy heavy-ion regime largely unexplored. We present, to our knowledge, the first particle-level foundation model pretrained directly on experimentally recorded heavy-ion collisions released through the CERN Open Data Portal. The training corpus is the CMS HIAllPhysics dataset from the 2010 PbPb run at (\sqrt{s_{\mathrm{NN}}}=2.76~\mathrm{TeV}), containing 75.5 million events. Each event is represented as a variable-length sequence of reconstructed particle-flow candidates. Self-supervised masked particle modelling is used to reconstruct masked kinematic features and particle categories without simulation-derived labels. Using a common particle interface and matched training protocol, we compare attention-based, state-space, hybrid, language-model-inspired, and transferred proton–proton collider architectures. We study scaling with training-set size, model size, and compute, and assess the learned representations through held-out masked reconstruction and three heavy-ion downstream tasks: particle identification, hard-jet versus combinatorial-jet discrimination, and jet-response correction. This work establishes an open and reproducible path from collision data to reusable heavy-ion event representations and provides a controlled test of which sequence-model inductive biases scale most effectively with heavy-ion multiplicity.