Speaker
Description
In recent years, several pre-training strategies have been proposed for foundation models in jet physics. These approaches range from self-supervised generative tasks like next-token prediction (NTP) and masked particle modeling (MPM), to standard supervised classification. Inputs to foundation models are often tokenized, but this leads to a loss of information. Recent work has shown that using a hybrid setup with continuous feature inputs can prevent this loss of information. However, finding the best way to combine these different pre-training objectives remains an open question. In this study, we investigate whether combining these diverse pre-training tasks translates to improved model efficacy on downstream applications. We systematically evaluate a variety of pre-training setups, including supervised classification alone, classification combined with MPM, classification with NTP, and a joint strategy utilizing all three objectives. We will highlight the general trade-offs between discriminative performance and generative capabilities to help guide the future development of robust collider physics foundation models.