Speakers
Description
The CMS 40 MHz Scouting program at the High-Luminosity LHC requires aggressive real-time compression of particle-flow event information to operate within stringent bandwidth constraints. We investigate the use of quantized autoencoders for converting Level-1 event representations into compact sequences of discrete tokens. We explore vector quantization, finite scalar quantization, and split-quantizer architectures that aim to separately encode morphological and amplitude information. Using the Collide-2V dataset, we evaluate these tokenization schemes in terms of compression rate, codebook utilization, constituent-level reconstruction fidelity, and preservation of jet observables including transverse momentum, mass, and substructure. We further evaluate the methods for downstream physics analyses by studying the reconstruction of the $H\rightarrow b\bar{b}$ dijet mass spectrum. Overall, these studies characterize the trade-off between token bitrate and physics fidelity, demonstrating the potential of learned tokenization as a physics-aware compression strategy for low-level particle data within the Scouting system.