14–18 Sept 2026
Europe/Vienna timezone

Scaling Laws for the Unified Particle Transformer in Heavy-Flavour Tagging

17 Sept 2026, 17:10
20m

Speaker

Pavlo Kashko (Vrije Universiteit Brussel (BE))

Description

Transformer-based taggers have become the workhorse of heavy-flavour identification at the LHC, with the Unified Particle Transformer (UParT) folding flavour classification, track-level auxiliary tasks and regression into a single architecture. Their development has nonetheless remained largely empirical: model capacity, training-sample size and input granularity are chosen by convention rather than derived from a predictive framework. We present a systematic study of the scaling behaviour of UParT-style taggers. Varying model size and training statistics over a wide range, we fit the tagging loss to the parameterization L(N, D) = E + A/N^α + B/D^β and observe clean power-law scaling in both quantities, together with a clearly non-zero irreducible term. We argue that this floor is physical rather than architectural and interpret it in terms of the achievable light-jet rejection at fixed b-tagging efficiency and other flavour tagging metrics.

Beyond the individual power laws, we study where each configuration saturates and extract the compute-optimal allocation between parameters and training jets, in direct analogy to the Chinchilla token-to-parameter analysis for language models. Strikingly, the optimal jet-to-parameter ratio follows a very similar trend, suggesting that the compute-optimal balance is governed by the structure of the training objective rather than by the specifics of the data domain. We further quantify how loss improvements propagate into misidentification rates and how the resulting Pareto front interacts with the inference-latency budgets of offline reconstruction and the high-level trigger.

Author

Pavlo Kashko (Vrije Universiteit Brussel (BE))

Presentation materials

There are no materials yet.