Speaker
Description
In collider-based particle physics experiments, proton collision events are commonly represented as tabular datasets for specific final states, features being four-vectors of final state particles or higher-level variables (invariant masses).
Inspired by the success of foundation models in language and vision, recent developments have introduced tabular foundation models such as TabPFN and TabICL.
We explore the use of tabular foundation models in collider physics analyses, and we show that they clearly outperform baselines like XGboost in data-limited regimes, for 50K training events or less. PFNs are benchmarked against boosted decision trees and neural networks across classification, density ratio evaluation (including calibration), and generation using the FAIR Universe dataset. We discuss performance, data efficiency, and computational trade-offs, and assess the potential role of pretrained tabular models in HEP analysis workflows.