Speaker
Description
Full simulation and reconstruction are projected to become major bottlenecks for computation at the High-Luminosity LHC, motivating the need for fast, ML-based surrogates. At the same time, LLMs have driven fast progress in generative discrete modeling: autoregressive transformers trained on tokenized data now represent the state of the art across a range of generative tasks. We extend the discrete modeling paradigm by introducing a general-purpose particle-level generative model trained on tokenized full-event data. We demonstrate the ability of this model family to perform both conditional generation from detector-stable particles and unconditional generation; study its scaling behavior across dataset and model sizes, and show that the learned representations transfer to downstream tasks.