ML4Jets 2026

Europe/Vienna
Claudius Krause (HEPHY Vienna (ÖAW))
Description

Over the past years, High Energy Physics (HEP) has made significant strides by developing and applying machine learning (ML) and artificial intelligence (AI) approaches. These advancements have greatly improved particle and event identification, reconstruction, simulation, experiments operations, and more.

The workshop will highlight the latest progress and ongoing challenges in these areas. It is open to the entire community, including participants from LHC experiments (detector and accelerator), theorists, and phenomenologists. We welcome contributions from method scientists and experts in related fields such as astronomy, astrophysics, cosmology, astroparticle physics, hadron and nuclear physics, and other domains facing similar challenges as well as computer scientists in research, industry and academia.

Join us to explore new ideas and advance the future of particle physics and science through ML/AI.

The following topics accross the various fields are foreseen:

  • Classification and reconstruction
  • Experiment simulation
  • Event generation
  • Inverse problems
  • Uncertainties
  • Anomaly detection
  • Interpretability
  • FastAI / EdgeAI systems for trigger applications
  • Detector/accelerator monitoring, control, and data acquisition
  • Agentic AI and LLMs to accelerate research procedures

 

Confirmed Plenary Speakers include:

  • Nicole Hartman
  • Benoit Assi
  • Theo Heimel
  • Sioni Summers
  • Eric Moreno
  • Anna Hallin
  • Marie Hein
  • Michal Mazurek
  • Jay Sandesara
  • Felix Weiglhofer
  • Manuel Szewc 

 

Registration is now open!  

For inquiries, please contact ml4jets2026@oeaw.ac.at

 


With the generous support of CERN's Next Generation Triggers Project and the Physics Faculty of the University of Vienna

MBI Vienna LogoÖAW LogoPhysics Faculty LogoUni Wien LogoNGT Logo

Registration
Registration
Participants
    • 08:30 09:00
      Registration 30m
    • 09:00 09:30
      Organisatoricals: Welcome
      Convener: Dr Claudius Krause (MBI Vienna (ÖAW))
    • 09:30 10:30
      Plenary Presentations
      • 09:30
        Experimental Overview 30m

        TBA

        Speaker: Nicole Michelle Hartman (TUM (DE))
      • 10:00
        Theory Overview 30m

        TBA

        Speakers: Ben Assi, Ben Assi, Benoit Assi
    • 10:30 11:00
      Coffee Break 30m
    • 11:00 12:00
      Parallel Talks
      • 11:00
        b-hive: a CMS wide Machine Learning Framework 20m

        b-hive is a general-purpose machine learning framework developed for the CMS experiment. Every stage of the workflow, from dataset construction and training to inference and evaluation, is encapsulated in a self-contained Law task, while physics-specific choices such as architecture, input features, truth definitions, and kinematic selections are injected through YAML configuration files and Python modules. The same core infrastructure therefore serves arbitrary classification and regression tasks on complex, variable-length HEP data. ROOT files are processed columnar-wise with coffea into LZ4-compressed datasets served by a custom iterable data loader; combined with mixed-precision training and torch.compile, this yields a factor 12.8 speed-up over the previous DeepJetCore-based setup, reducing a week-long training to roughly 13 hours. b-hive ships implementations of model jet tagging algorithms, together with a fully configurable adversarial module for robustness studies and adversarial training. The framework was used to develop the Unified Particle Transformer v2, the official CMS jet tagging algorithm for Run 3. This presentation covers the design, current performance, and planned developments.

        Speaker: Ulrich Willemsen (Rheinisch Westfaelische Tech. Hoch. (DE))
      • 11:20
        Machine Learning Tools for JetMET Data Certification in CMS 20m

        The CMS experiment at CERN relies on Data Quality Monitoring (DQM) and data certification to ensure that only high-quality data are used for physics analyses. For the JetMET subsystem, this process is traditionally based on the manual inspection of a large number of DQM histograms by detector experts, making it a time-consuming task that can make subtle detector or reconstruction issues difficult to identify consistently. To support the JetMET offline certification workflow, we first explored unsupervised anomaly detection using autoencoders to identify runs with anomalous detector or reconstruction behaviour from DQM histograms. Building on this work, we integrated machine learning models into the Data Inspector for Anomalous Lumi-Sections (DIALS), a framework providing access to per-lumisection DQM information, and developed a workflow to produce Machine Learning JSON (MLJSON) certification files containing the run and lumisection selections for physics analyses. We also developed a web application to simplify model training, validation, and result visualization for JetMET experts. This contribution presents the machine learning developments introduced for JetMET during Run 3, their integration into the CMS certification workflow, and the ongoing work towards preparing these tools for the High-Luminosity LHC era.

        Speaker: Jawaher Altork (Universita e INFN, Firenze (IT))
      • 11:40
        Improving particle identification in the Belle II TOP detector using machine learning 20m

        The Time Of Propagation (TOP) detector at the Belle II experiment is a ring-imaging Cherenkov detector designed to identify charged hadrons in electron-positron collisions at the SuperKEKB accelerator. It consists of 16 quartz radiator modules arranged around the barrel region of the Belle II detector. When a charged particle crosses a module, Cherenkov photons are emitted. A fraction of these photons is trapped by total internal reflection and propagates through the quartz to an array of micro-channel-plate photomultiplier tubes (MCP-PMTs), which measure the arrival time and hit position of the photons. Since the Cherenkov angle depends on the particle species for a given momentum and direction, different particles produce distinct position-time photon patterns. Particle identification (PID) in the TOP detector therefore reduces to a pattern recognition problem.
        Currently, PID in the TOP relies on comparing event-by-event photon patterns with analytical probability density functions (PDFs) computed for different particle hypotheses and track parameters. So far, TOP PID performance has been consistent with its design expectations. However, further improvements are limited by the need for highly accurate modeling of the detector geometry and optical properties.
        Machine Learning (ML) techniques offer a data-driven approach for PID in the TOP by learning the relationship between photon patterns and particle identity directly from data. In particular, Convolutional Neural Networks (CNNs), which are specifically designed for pattern recognition tasks, are well suited to analyze the two-dimensional position-time photon patterns recorded by the TOP detector. This approach has the potential to improve PID performance while reducing sensitivity to residual detector mismodeling.
        A feasibility study of a CNN-based approach has been conducted using simulated samples of charged pions and kaons divided into phase-space bins according to the track momentum and direction. The results, validated on both simulated and experimental data, show an improvement in the kaon/pion separation performance of the TOP detector with respect to the standard PDF-based likelihood method. However, the performance gain is achieved at the cost of a considerably more complex PID approach that requires significantly greater computational resources.

        Speaker: Cecilia Antonioli (INFN & Universita' di Padova)
    • 11:00 12:00
      Parallel Talks: Uncertainty Quantification
      • 11:00
        Model Uncertainty in the Measurement of the Gluon-Jet Fraction 20m

        The measurement of quark- and gluon-jet properties requires the determination of their fractions in experimental jet samples. This is commonly achieved by fitting data with quark- and gluon-jet templates derived from Monte Carlo simulations. Because the template shapes depend on the underlying model, the extracted jet fractions are intrinsically model-dependent. This talk reviews the main sources of model uncertainty in gluon-jet fraction measurements, their impact on the results, and methods for their evaluation in LHC data analyses.

        Speaker: Siarhei Shulha (Joint Institute for Nuclear Research (RU))
      • 11:20
        Local Conformal Predictions for Calibrated Surrogates 20m

        Neural network surrogates for LHC scattering amplitudes require trustworthy uncertainty estimates, a challenging task given the non-Gaussian systematics. We target it
        using conformal prediction, a distribution-free post-processing to complement trained
        surrogates with calibrated uncertainties. We find that standard conformal predictions
        struggle to provide locally calibrated uncertainties. This leads us to introduce FALCON,
        a novel conformal prediction method that learns locally calibrated confidence intervals.
        Our simple examples illustrate the power of distribution-free uncertainty quantification
        for ultra-fast event generation at the LHC.

        Speaker: Suprio Dubey (Heidelberg University)
      • 11:40
        Know What You Don't Flow 20m

        Calibrated learned uncertainties are a key requirement also for generative neural networks in LHC physics. For a toy model with an explicit likelihood we show how a heteroscedastic and a Bayesian normalizing flow learn the systematic and statistical uncertainties on the underlying phase space density. Without an explicit likelihood we train the heteroscedastic loss on a classifier-reweighted approximate generative network. We illustrate our comprehensive approach for top pair events and show how a conditional heteroscedastic flow propagates calibrated uncertainties to all phase space directions.

        Speaker: Lorenz Vogel (Institute for Theoretical Physics, Heidelberg University)
    • 12:00 13:30
      Lunch Break 1h 30m
    • 13:30 15:30
      Parallel Talks: Jet Tagging
      • 13:30
        Transforming Flavour Tagging with the ATLAS Detector: ML tools and calibration techniques for small-R jets and boosted objects 20m

        The identification of jets containing b-hadrons is essential for many physics analyses at the LHC, including precision measurements of Higgs boson and top-quark processes, as well as searches for physics beyond the Standard Model. We present recent improvements in the discrimination of b-jets from jets originating from lighter quarks using the ATLAS detector. These advances are driven by state-of-the-art machine learning techniques based on transformer architectures. Their performance is well modelled by the ATLAS simulation, as demonstrated through dedicated calibration studies, the results of which will be presented. Compared to previous algorithms, the transformer-based approach improves the rejection of c-jets (light-jets) by factors of 3.5 (1.8) at a b-jet tagging efficiency of 70%. We also discuss the latest version of this algorithm and its expected performance at the High-Luminosity LHC (HL-LHC). Recent advances in Higgs-boson identification algorithms for cases where the Higgs-boson decay products are captured in a single large-radius are presented as well. This algorithm improves the inclusive QCD jet rejection by approximately 25% and the rejection of fully-contained top-jets by a factor of two for a the H->bb tagging efficiency of 70%.

        Speaker: Alexander Gavin (University of London (GB))
      • 13:50
        Deep Learning Methods for Jet Tagging and Process Classification Using Image Processing 20m

        This study explores a convolutional neural network (CNN) approach to classify events produced in high-energy collisions by the presence of heavy (charm and bottom), light (up, down, strange) and gluon jets, with the main characteristic being that jets are not reconstructed in our approach. The method constructs image-like representations based on the kinematics of charged decay products using detector-level variables, which allow CNNs to identify visual patterns characteristic of each jet type. As an ongoing step, we also study model sensitivity to jet properties in reconstructed jets with the intent to translate them to an event landscape that can increase our models performance. This approach not only demonstrates strong classification performance, highlighting the versatility of CNN architectures in jet tagging, but also reveals the abillity of AI methods to recognize structures that can be associated to flavor specific jets, even in the absence of jet reconstruction.

        Speaker: Jhoão Gabriel Martins Campos de Almeida Arneiro (Universidade de São Paulo (USP))
      • 14:10
        An ML-based 4-prong Tagger for hadronic decays of highly boosted heavy scalars in ATLAS 20m

        Discriminating highly boosted jets from the decays of heavy scalar particles that do not involve b-quarks from background is a very challenging problem. We describe a novel "4-Prong Tagger" that uses a graph neural network based on the Lund Jet Planes of a large radius jet to distinguish high energy scalar particle decays, S->WW->4q, from backgrounds such as QCD, hadronic vector-boson decays, and hadronic top decays. We show that this tagger can provide substantial sensitivity gains with respect to other jet sub-structure-based discriminants using a simple example analysis. We also discuss techniques for calibrating the 4-Prong Tagger using the Lund Jet Planes of the subjets of the top and vector boson control samples.

        Speaker: Jae Jin Hong (Indiana University (US))
      • 14:30
        AI for Rare Top-Higgs Processes at the LHC 20m

        The observation of flavor-changing neutral current (FCNC) interactions between the top quark and the Standard Model (SM) Higgs boson would constitute an unambiguous signal of physics beyond the SM. Searches for this process at the LHC are, however, extremely challenging due to the small signal rates and the strong kinematic resemblance between the signal and dominant SM backgrounds, particularly top-quark pair production accompanied by QCD jets. In this talk, I will present a novel deep-learning approach that exploits event-level properties such as QCD color flow through graph-based neural network architectures. By comparing its performance with more traditional methods, including multilayer perceptrons (MLPs), I will demonstrate the potential of these techniques for integration into ATLAS and CMS analysis frameworks, enhancing the sensitivity to this channel.

        Speaker: Dr Adil Jueid (Korea Institute for Advanced Study)
      • 14:50
        Simplex Demixing: Disentangling Multiple Light-Flavor Jets at Colliders 20m

        Providing a practical and hadron-level definition of multiple jet flavors has been a long-standing challenge in collider physics. Previous work has introduced a data-driven, operational definition of quark and gluon jets, but no robust generalization beyond two jet categories presently exists. To address this, we introduce a machine-learning framework called "simplex demixing'' to extract $T$ jet flavors (or topics in the statistics literature) from $M$ data samples (or mixtures) with minimal constraints. Intuitively, our procedure identifies the maximally separable categories in the data, translating a multi-category classifier on the $M$ mixtures into a bounded geometric object with $T$ vertices. We first demonstrate our procedure on a toy problem to infer the truth-level fractions of down-quark, up-quark, and gluon jets from synthetic mixtures of the three pure samples. We then propose a tag-and-probe strategy to extract multiple light-flavor categories in a more realistic collider setting involving dijet production. As expected, the identifiability of jet flavors depends on their relative abundance in the samples and the hadron-level information available to the classifier architecture. Our work opens the door to data-driven extractions of multiple jet flavor properties at the Large Hadron Collider.

        Speaker: Gregorio de la Fuente Simarro (Massachusetts Institute of Technology)
      • 15:10
        "Hadron-in-fat-jet'' AI Tagging to Detect Rare Decays 20m

        We investigate a novel class of boosted-object signatures at the LHC, where a high-pT fat-jet contains an identifiable hadron or quarkonium state originating from rare or semi-exclusive decays. Unlike conventional boosted jet studies, which focus on multi-prong partonic substructure, our approach probes hybrid configurations such as W±→π±γ, where a localized hadronic or quarkonium signal is embedded within a collimated jet. By fine-tuning the signature-oriented, pre-trained Sophon AI model optimized for large-radius jets, and combining it with an event-level BDT and a soft-drop-mass shape fit, we obtain an expected 95\% CL upper limit of (W±→π±γ)<2.78×10−5 for 450fb−1 in our nominal setup. This study serves as a first proof-of-principle demonstration of the ``hadron-in-fat-jet'' paradigm; substantial gains in sensitivity are expected from improved trigger strategies, additional production channels, and dedicated taggers, while the methodology itself is broadly applicable to a wide range of rare Standard Model processes and searches for light or exotic resonances at present and future collider experiments. https://arxiv.org/abs/2606.09458

        Speakers: Mr Linrui Chen (Peking University), Qiang Li (Peking University (CN)), Mr Zixun Kou (Peking University)
    • 13:30 15:30
      Parallel Talks: Phenomenology
      • 13:30
        Mass-unspecific classifiers for mass-dependent searches 20m

        Searches for new particles often span a wide mass range, where both signal and SM background shapes vary significantly. We introduce a multivariate method that fully exploits the correlation between signal and background features and the explored mass scale. The classifiers—either a neural network or boosted decision tree—produce continuous outputs across the full mass range, achieving performance similar to classifiers trained for the specific mass.

        The key advantages arise from two factors:

        1. The background mass scale is correlated with the actual background shape, enabling more effective background identification across all mass scales.

        2. A balanced training sample that spans the entire mass range, allowing the classifier to learn the differences between high and low scales.

        We benchmark this approach with single production of a vector-like quark singlet T at the HL-LHC, where the cross section depends on both the mixing angle and the quark mass. Our method is effective for mass-unspecific searches, applicable to a wide range of new physics processes and collider settings. Mass-unspecific classifiers show strong performance, especially in searches spanning a broad mass range.

        Speaker: SERGIO RODRIGUEZ BENITEZ (Instituto de Física Teórica IFT-UAM/CSIC)
      • 13:50
        A Machine-Learning Analysis of the Higgs Boson in the H → ZZ∗ →4ℓ Channel 20m

        The H → ZZ → 4ℓ channel remains one of the cleanest probes of Higgs boson properties at the LHC, owing to its fully reconstructable final state and well-understood background composition. This work addresses the problem of identifying an optimal machine-learning classifier for extracting the H → ZZ → 4ℓ signal from the ATLAS Open Data 2025 release (√s = 13 TeV, 36.6 fb⁻¹), comparing seven algorithms, XGBoost, LightGBM, Random Forest, a multilayer perceptron, Logistic Regression, QDA, and Gaussian Naive Bayes, under a single, controlled experimental protocol. The specialty of this analysis is twofold. First, every classifier is trained exclusively on Monte Carlo simulation and then applied, without retraining or recalibration, to the real ATLAS dataset, so that the reported significance reflects genuine generalization from simulation to data rather than performance on held-out simulated events alone. Second, background estimation itself is treated as a robustness test: each classifier's signal significance on real data is computed twice, once under a Monte Carlo background prediction and once under a data-driven sideband extrapolation, and a classifier is judged reliable only if the two estimates agree within uncertainties, since the two methods rest on largely independent assumptions. Four-lepton events are reconstructed from the ATLAS Open Data samples and reduced to a set of 31 physics-motivated kinematic and angular features, subsequently pruned to 20 via separation-power ranking and correlation filtering. All seven classifiers are trained and cross-validated on Monte Carlo simulation under a 5-fold stratified scheme, then applied at a fixed classifier-score threshold to the full real-data sample, with the signal region and sideband control regions defined directly in the four-lepton invariant mass spectrum. Signal significance is computed independently under both background estimators for each classifier, yielding a direct test of consistency rather than a single, isolated figure of merit. Under this framework, LightGBM emerges as the best-performing and most consistent classifier overall, achieving a signal significance of Z = 4.66 ± 1.22σ (p = 1.55 × 10⁻⁶) under the Monte Carlo background estimate and Z = 5.08 ± 1.50σ (p = 1.90 × 10⁻⁷) under the sideband estimate. XGBoost follows closely, with Z = 4.24 ± 1.10σ (p = 1.14 × 10⁻⁵) and Z = 4.70 ± 1.35σ (p = 1.31 × 10⁻⁶) under the same two methods, respectively. Both classifiers reach evidence-level significance under both background estimators, with their results agreeing within uncertainties across two independent methods, confirming that the models generalize from simulation to real collision data and that the resulting evidence is not an artifact of a particular background model. This consistency establishes gradient-boosted decision trees as a robust, reproducible means of recovering Higgs boson evidence from public ATLAS Open Data using standard machine-learning methods alone.

        Speaker: Mane Papoyan (American University of Armenia (AUA))
      • 14:10
        Hunting the Unseen: Parameter-Agnostic Deep Learning for Semi Visible Jet Tagging 20m

        In the study of DM detection at colliders, novel candidates have emerged to bridge between experimental data and theoretical models. Dark Showers (DS) are being studied as an extension of the Standard Model (SM), containing both invisible and visible particles that allow us to predict scenarios involving collider observables. Among these, Semi-Visible Jets (SVJs), represent a novel promising new signature, particularly those mediated by a massive $Z^\prime$ boson that enables the production of heavy dark hadrons. In such a scenario, jets enclose visible SM particles and invisible dark matter components, having more Missing Transverse Energy $E_{T}^{miss}$ (MET). The complexity is further heightened since the poorly constrained nature of the dark sector; factors such as the dark hadronisation constant, dark meson masses and the fraction of the invisible DM hadrons will be part on the extense amount of hidden sector parameters that manage how the standard model hadrons will decay eventually into kinematical observables. Our study will review different tools at a 2-component analysis, global-level jet kinematics and jet-substructure level to study Semi-Visible Jets (SVJs), formed by the decay of a resonant heavy gauge boson $Z^\prime$. We characterise a di-jet system focused on final-state global jet observables such as the transverse momentum, azimuthal angle and pseudorapidity $(p_{T},\phi,\eta)$, with modern related MET studies as the global MET and the difference on the azimuthal angle between the leading jet and the global MET, alonside jet-substructure metrics such as the Energy-Energy Correlation Functions (EECs), Angularity ($\tau$), SM Hadrons multiplicity and the Lund Jet Plane (LJP) as kinematical features. We do all this characterisation since our ultimately goal is to differentiate between QCD and SVJs event signals with modern Deep Learning (DL) analyses, we implemented a Vision Transformer ($\text{ViT}$) neural network (NN) focused on Lund plane images and a MultiLayer Perceptron ($\text{MLP}$) NN for the rest of high-level observables and finally use evaluation metrics like $\text{AUC}$, accuracy and $\text{ROC}$ curve to evaluate not only the classification performance but also the signal rejection in SVJs searches for future experimental searches.
        It is worthly saying that the paper should be soon on arxiv.

        Speaker: Mr Miguel Angel Avendano Bernal (University of Southampton)
      • 14:30
        Understanding the Performance Gains of a full-event HH→4b Search 20m

        A calibratable full-event (jet-free) HH→4b framework based on full-event particle-flow (PF) candidates was presented at ML4Jets 2025 and in Refs. 1 and 2, demonstrating a dramatic >5× improvement in search sensitivity over conventional approaches. In this talk, I will present a series of controlled ablation studies that progressively enhance the event representation, providing a detailed understanding of the origin of this dramatic gain in background rejection. I will also introduce an improved training strategy for mass-decorrelated discriminants and a dedicated optimization targeting constraints on the Higgs self-coupling, $\kappa_{\lambda}$.

        Starting from a conventional jet-based baseline, in which the kinematic variables and flavour-tagging scores of all small-radius jets are used to train a signal-versus-background event classifier, we progressively enhance the event representation through a sequence of modifications: (a) replacing the handcrafted jet features with 64-dimensional learned jet embeddings extracted from a jet-tagging network (~2× gain in background rejection); (b) jointly training the particle-to-jet encoder (i.e. the jet-tagging network) with the event-level classifier, making the jet representations task-specific (~1.5× gain); (c) removing the per-jet information bottleneck and directly modeling particle-level correlations across jets, yielding the largest individual improvement (~3×); (d) incorporating PF candidates outside reconstructed jets (~1.5–2× gain); and (e) scaling up both the training statistics and model capacity (~2× gain). Altogether, the full-event PF representation achieves ~20× stronger background rejection than the conventional jet-based baseline at representative working points.

        Finally, we extend the framework by jointly training on the $\kappa_{\lambda} = 0, 1, 2.45,$ and $5$ hypotheses, allowing the model to learn $\kappa_{\lambda}$-dependent event features and construct a discriminant specifically optimized for constraining the Higgs self-coupling. Preliminary results indicate strong potential for excluding the $\kappa_{\lambda} = 0$ hypothesis.

        Speaker: Yipin Wang (Peking University (CN))
      • 14:50
        Enhancing the Sensitivity for Triple Higgs Boson Searches with Deep Learning Techniques 20m

        Using two benchmark models containing extended scalar sectors beyond the Standard Model, we investigate deep learning techniques to enhance the sensitivity of resonant triple Higgs boson ($HHH$) searches in the fully hadronic $6b$ channel, which suffers from severe combinatorial background and jet-pairing ambiguities. Specifically, we employ the Symmetry Preserving Attention Network (SPA-Net), a Transformer-based architecture that explicitly respects the permutational symmetries inherent in jet assignment. By performing multi-task learning to simultaneously tackle jet pairing and event classification directly from low-level jet features, SPA-Net eliminates the need for explicit jet permutation enumeration while building expressive latent event representations. Compared with conventional Dense Neural Networks, SPA-Net yields up to 40% more stringent limits on resonant production cross-sections. These results highlight the potential of symmetry-preserving deep learning models to overcome combinatorial barriers in high-multiplicity hadronic searches.

        Speaker: Feng-Yang Hsieh (National Taiwan University)
      • 15:10
        Measurement of the Energy Dependence of Strong Isospin Violation in $\Upsilon(4S)$ Decays Using a Neural Network Event Classifier 20m

        Precise measurements of $B$ meson branching fractions are essential both for testing Standard Model predictions and for many measurements that rely on accurate modeling of data composition in flavor-physics analyses. At Belle II, $B$ mesons are produced via the process $e^+e^- \to \Upsilon(4S) \to B\bar{B}$. The total number of such events can be determined with high precision, yet using this quantity for $B$ meson branching fraction measurements requires knowledge of the production fractions of neutral and charged B meson pairs in $\Upsilon(4S)$ decays, respectively $f_{00}$ and $f_{\pm}$. Although strong isospin symmetry predicts equal production rates, recent theoretical studies indicate a possible energy dependence of the ratio $R^{\pm0}=f_{\pm}/f_{00}$, motivating an experimental investigation. Reconstructing specific $B$ decay channels restricts the analysis to a small fraction of the available $B\bar{B}$ dataset, motivating an inclusive, machine-learning-based approach that exploits the full dataset without channel-specific assumptions.
        We present a measurement of $R^{\pm0}$ using a neural network classifier trained on event-level observables to separate $B^+B^-$ from $B^0\bar{B}^0$ events without reconstructing either $B$ meson explicitly. A simultaneous three-class classifier and a two-stage binary architecture, which first suppresses background before separating the two $B \bar{B}$ event types, were both explored to address the limitations of the classification task. $R^{\pm0}$ is extracted from the classifier output through a binned template fit. We discuss the classifier design, the systematic effects relevant to its application on data, and the outlook for extracting the energy dependence of $R^{\pm0}$ from Belle II data.

        Speaker: Lena Nowatzki
    • 15:30 16:00
      Coffee Break 30m
    • 16:00 18:00
      Parallel Talks: Jet Tagging
      • 16:00
        Instance-Level b-Hadron Tagging with DELPHI Open Data 20m

        The release of LEP open data provides a clean test bed for developing machine-learning methods relevant to future lepton colliders. We present an effort to identify individual b hadrons in DELPHI $Z\to b\bar{b}$ events using an approach inspired by panoptic segmentation. Rather than assigning a single flavour label to a jet, the model classifies reconstructed particles and associates them with the decay products of each b hadron. We propose a design of conversion to EDM4HEP format for DELPHI open-data, the formulation of particle-to-hadron assignment as a segmentation problem, and initial tagging studies. This work explores archival DELPHI data as a benchmark for instance-level heavy-flavour reconstruction.

        Speaker: Sitian Qian (NU FNAL)
      • 16:20
        Domain Adaptation in Colliders: Reconcile Simulation and Data 20m

        Multivariate classifiers in collider physics increasingly use low-level detector information to maximize sensitivity, but this often amplifies systematic uncertainties arising from mismodelling in simulation. To address this, we apply unsupervised domain adaptation (UDA), allowing the classifier to learn with both simulated and unlabeled real data during training, thereby reducing sensitivity to simulation-induced domain shifts. Our approach leverages a maximum mean discrepancy (MMD) penalty as a plug-in module for standard deep learning architectures. Using a case study of VBF versus ggF Higgs classification in H→γγ with both CNN and Particle Transformer architectures, we show that this method consistently improves classifier robustness and reduces cross section measurement uncertainty, demonstrating its universal effectiveness across data types.

        Speaker: Shang-Fu Wei (The University of Tokyo)
      • 16:40
        Classifying hadronic objects in ATLAS with machine learning 20m

        Hadronic object reconstruction & classification is one of the most promising settings for cutting-edge machine learning and artificial intelligence algorithms at the LHC. In this contribution, recent highlights of ML applications by ATLAS for boosted-object identification will be presented. This covers results of constituent-based transformers for quark-gluon, top-quark, and W boson discrimination, as well as detailed studies in performance of multi-class taggers, MC generator dependencies, and polarization-aware tagging.

        Speaker: Robert Les (Michigan State University (US))
      • 17:00
        PanopTag: Simultaneously Tagging All Jets in a Particle Collision Event 20m

        Jet tagging, identifying the origin of jets produced in particle collisions, is a critical classification task in high-energy physics. Despite the revolutionary impact of deep learning on jet tagging over the past decade, the paradigm has remained unchanged. In particular, jets within the same collision event are classified independently, one at a time. This single-jet approach ignores correlations, overlaps, and wider event context between jets that can greatly improve classification performance. We introduce PanopTag, a new paradigm for jet tagging that simultaneously tags all jets by employing an encoder-decoder architecture that uses jet kinematics as queries to cross-attend to particle flow object embeddings from the full collision event. We evaluate PanopTag on heavy-flavor $(b/c)$-tagging and demonstrate remarkable performance improvements over state-of-the-art single-jet baselines that are only accessible by exploiting event-level
        features and correlations between jets. We demonstrate the absence of topology bias and conditional independence of jet flavor predictions, showing that PanopTag could utilize established calibration strategies for ultimate deployment in collider data analysis.

        Speaker: Umar Sohail Qureshi (Vanderbilt University)
      • 17:20
        Physics Aware Modifications to Particle Transformers (Mod-ParT) for Jet Tagging 20m

        Particle Transformer ( ParT) has achieved state of the art performance in jet tagging by modeling particle level information through attention. In this work, we introduce Mod-ParT, a physics aware extension of ParT that incorporates relational information into particle representations, enriches pairwise attention biases with additional physics motivated features and replaces the class token and class attention stages with a hybrid pooling module for jet representation. We trained Mod-ParT from scratch on the established top tagging reference dataset and the large scale TopTagXL dataset, using training samples of 1.2 million and 10 million jets respectively. We further trained the model on quark gluon discrimination benchmark, generated with Pythia8. Across the top tagging and quark gluon discrimination tasks, Mod-ParT achieves competitive performance in terms of accuracy, AUC and background rejection metrics. We also investigate computational performance under matched execution conditions on a single GPU considering computational complexity, training latency and peak memory consumption. The results show that Mod-ParT can substantially reduce computational and memory requirements relative to ParT and several architectures evaluated under comparable benchmark conditions while maintaining competitive tagging performance. Overall, Mod-ParT provides a favorable balance between predictive performance and lower computational cost and higher speed across distinct jet classification tasks.

        Speaker: Mr Osman Bayraktar (Istanbul University (TR))
      • 17:40
        Multiclass classification without labels (multiCWoLa) 20m

        In many classification problems, reliable instance-level labels are unavailable. However, it is often possible to construct weakly enriched unlabeled samples: datasets selected by different cuts, sources, populations, or experimental conditions that change latent class proportions without revealing them. Classification without Labels (CWoLa) shows that, in the binary case ($K=2$), a classifier trained to distinguish two impure mixtures with different class proportions can recover an optimal class discriminator without knowing the mixture proportions. We extend this principle to multiclass learning from several unlabeled mixtures ($K>2$), where the learner observes only mixture identity and neither latent class labels nor class-prior matrices. We prove that, for a multiclass mixture model, the Bayes-optimal mixture classifier $g^\star$ maps data points into a $(K-1)$-simplex embedded in mixture-posterior space. The $K$ vertices of this simplex are induced by the latent classes through the unknown mixing matrix. Leveraging this geometry, we propose prior-free procedures that train a standard classifier to distinguish mixture identities and then extract latent class structure using either post-hoc simplex fitting or a bottleneck architecture. Experiments on MNIST, CIFAR-10, and Galaxy10 DECaLS show that mixture identity alone can recover latent classes and their fractions in the mixture. By narrowing the gap between weakly supervised and fully supervised performance, we provide a mathematically grounded, scalable tool for multiclass discovery in label-scarce domains.
        In this talk, we also present applications to particle physics and astrophysics.

        Speaker: Johann Michael Ioannou-Nikolaides (University of Copenhagen (DK))
    • 16:00 18:00
      Parallel Talks: Reinterpretation & Theory
      • 16:00
        Clustering for large-scale reinterpretation of new physics searches at the LHC 20m

        Beyond the Standard Model (BSM) searches show no statistically significant sign of new physics to date. However, several analyses reported small excesses, higher than 2σ SD beyond the SM expectation. In this work, clustering algorithms are used to extract more insight from existing searches and motivate a next round of BSM analyses. The flexible framework of the phenomenological Minimal Supersymmetric Standard Model (pMSSM) is used for a large-scale reinterpretation of new physics models using existing analyses. Models consistent with the observed excesses are explored using clustering methods. In this work, clustering algorithms are benchmarked for this high-dimensional problem. Optimal algorithms are deployed to perform a large-scale reinterpretation of new physics searches to help inform the direction of future searches at the HL-LHC.

        Speaker: Judita Mamuzic (IFAE - Barcelona)
      • 16:20
        Open library of learned likelihoods from LHC experiments 20m

        Open statistical models released by the LHC experiments are transforming the reinterpretation of collider searches by providing access to the full likelihood information of experimental analyses. However, evaluating these likelihoods remains computationally expensive, particularly for analyses with many signal and control regions, limiting their use in large-scale phenomenological studies. In this contribution, we present the Open Library of Learned Likelihoods (OLLL), a collection of neural-network surrogates for profiled likelihoods from public LHC statistical models. The surrogate models are trained on likelihood evaluations generated with pyhf, optimized using a heteroscedastic loss that enables them to predict both the profiled likelihood and an associated uncertainty estimate, and serialized in the framework-independent ONNX format for seamless integration into reinterpretation frameworks. We demonstrate the approach on five ATLAS SUSY searches of increasing complexity, where the learned likelihoods accurately reproduce both the official ATLAS and full pyhf exclusion contours while reducing the computational cost of likelihood evaluation by orders of magnitude. The resulting library provides a practical and extensible infrastructure for fast and statistically faithful reinterpretation of LHC searches.

        Speaker: Humberto Reyes-Gonzalez (RWTH Aachen University)
      • 16:40
        Learning to bin: differentiable approach for multi-dimensional discriminants in high-energy physics 20m

        Categorizing events using discriminant observables is central to many high-energy physics analyses. Yet, bin boundaries are often chosen manually. A simple, popular choice in multi-classification tasks is to assign events according to the largest per-class score (”argmax”) and to apply equidistant binning to the resulting one-dimensional discriminants. We propose a binning optimization for signal significance directly in multi-dimensional discriminants using a differentiable approach. We use a Gaussian Mixture Model (GMM) to define flexible regions in the score space, which can be interpreted either as bins or as analysis categories. While this GMM-based strategy is applicable in both one and multiple dimensions, we also study a direct bin-boundary optimization in one dimension as a simpler alternative for binary discriminants.

        This talk presents the methods and results of the paper arXiv:2601.07756, with a particular focus on the differentiable optimization approach. We evaluate the method on binary and multi-class toy classification problems, as well as on the FAIR Universe $H\rightarrow\tau\tau$ benchmark dataset, demonstrating its applicability in realistic analysis scenarios. The proposed approach is compared with the equidistant binning strategy and with alternative binning methods based on k-means clustering and Bayesian optimization.

        Speaker: Nitish Kumar Kasaraguppe Veerappa Gowda (Rheinisch Westfaelische Tech. Hoch. (DE))
      • 17:00
        Resummed Distribution Functions: Making Perturbation Theory Positive and Normalized 20m

        Fixed-order perturbative calculations for differential cross sections can suffer from non-physical artifacts: they can be non-positive, non-normalizable, and non-finite, none of which occur in experimental measurements. We propose a framework, the Resummed Distribution Function (RDF), that, given a perturbative calculation for an observable to some finite order in $\alpha_S$, will ``resum'' the expression in a way that is guaranteed to match the original expression order-by-order and be positive, normalized, and finite. Moreover, our ansatz parameterizes all possible finite, positive, and normalized completions consistent with the original fixed-order expression, which can include N$^n$LL resummed expressions. The RDF also enables a more direct notion of perturbative uncertainties, as we can directly vary higher-order parameters and treat them as nuisance parameters. We demonstrate the power of the RDF ansatz by matching to thrust to $\mathcal{O}$($\alpha^3_S$) and extracting $\alpha_S$ with perturbative uncertainties by fitting the RDF to ALEPH data.

        Speaker: Radha Mastandrea
      • 17:20
        Theory-informed neural networks for particle physics 20m

        We present a theory-informed reinforcement-learning framework that recasts the combinatorial assignment of final-state particles in hadron collider events as a Markov decision process. A transformer-based deep Q-network, rewarded at each step by the logarithmic change in the magnitude of the tree-level matrix element, learns to map final-state particles to partons. Because the reward derives solely from first-principles theory, the resulting policy is label-free and fully interpretable, allowing every particle to be traced to a definite partonic origin. The method is validated on event reconstruction for
        $t\bar{t}$, $t\bar{t}W$, and $t\bar{t}t\bar{t}$ processes at the large hadron collider. At this stage the method is tested on parton-level data only, with tests using momentum smearing used as a proxy for a full simulation. The method maintains robust performance across all processes, demonstrating its scaling with increasing combinatorial complexity. We demonstrate how this method can be used to build a theory-informed classifier for effective discrimination of longitudinal $W^+W^-$ pairs, and show that we can construct theory-informed anomaly-detection tools using background process matrix elements. Building on theoretical calculations, this method offers a transparent alternative to black-box classifiers. Being built on the matrix element, the classification and anomaly scores naturally respect all physical symmetries and are much less susceptible to the implicit biases common to other methods. Thus, it provides a framework for precision measurements, hypothesis testing, and anomaly searches at the high-luminosity LHC.

        Speaker: Barry Dillon (b.dillon@ulster.ac.uk)
      • 17:40
        Scaling Particle-Level Foundation Models on Real Heavy-Ion Collision Data 20m

        Foundation models for collider physics have so far been trained predominantly on simulated events or proton–proton collision data, leaving their behaviour in the high-occupancy heavy-ion regime largely unexplored. We present, to our knowledge, the first particle-level foundation model pretrained directly on experimentally recorded heavy-ion collisions released through the CERN Open Data Portal. The training corpus is the CMS HIAllPhysics dataset from the 2010 PbPb run at (\sqrt{s_{\mathrm{NN}}}=2.76~\mathrm{TeV}), containing 75.5 million events. Each event is represented as a variable-length sequence of reconstructed particle-flow candidates. Self-supervised masked particle modelling is used to reconstruct masked kinematic features and particle categories without simulation-derived labels. Using a common particle interface and matched training protocol, we compare attention-based, state-space, hybrid, language-model-inspired, and transferred proton–proton collider architectures. We study scaling with training-set size, model size, and compute, and assess the learned representations through held-out masked reconstruction and three heavy-ion downstream tasks: particle identification, hard-jet versus combinatorial-jet discrimination, and jet-response correction. This work establishes an open and reproducible path from collision data to reusable heavy-ion event representations and provides a controlled test of which sequence-model inductive biases scale most effectively with heavy-ion multiplicity.

        Speaker: João A. Gonçalves (University of Bonn, B-IT, Lamarr Institute)
    • 18:00 21:00
      Welcome Reception 3h
    • 08:30 09:00
      Registration 30m
    • 09:00 10:30
      Plenary Presentations
      • 09:00
        Event Generation Overview 30m

        TBA

        Speaker: Theo Heimel (UCLouvain)
      • 09:30
        Overview of the Application of Generative Models to Fast Simulation in Large HEP Experiments 30m

        Monte Carlo simulations are essential for physics analyses and detector design in High Energy Physics (HEP). As the computational requirements of traditional simulation rapidly outpace available computing budgets, leveraging Generative AI for fast simulation has become vital to producing the required volume of simulated samples.
        Nevertheless, transitioning from initial, toy models to full-scale production remains a major challenge. This review highlights the current landscape and operational deployment of generative fast simulation across major HEP experiments. We will evaluate leading architectural paradigms and address practical solutions, including framework integration, validation, and key lessons learned when moving from prototype to full-scale production.

        Speaker: Michał Mazurek (National Centre for Nuclear Research (PL))
      • 10:00
        SBI Overview 30m

        TBA

        Speaker: Jay Ajitbhai Sandesara
    • 10:30 11:00
      Coffee Break 30m
    • 11:00 12:00
      Parallel Talks: Generative Models
      • 11:00
        Generative Amplification with Surrogate Monte Carlo 20m

        Amplitude surrogates speed up one of the most computationally expensive steps in the LHC simulation chain. We adapt the generative amplification framework to amplitude surrogates evaluated by reweighting and apply it to uncertainty-aware surrogates. We show that amplitude surrogates, like generative event generators, can exhibit generative amplification when trained on a limited set of expensive amplitude evaluations. In particular, we study gluon-associated $Z$ production and find that amplification emerges in the sparsely populated high-$p_T^Z$ tails, where it matters most.

        Speaker: Rebecca Revelli (Heidelberg University)
      • 11:20
        Forecasting Generative Amplification 20m

        Generative networks are perfect tools to enhance the speed and precision of LHC simulations. It is important to understand their statistical precision, especially when generating events beyond the size of the training dataset. We present two complementary methods to estimate the amplification factor without large holdout datasets. Averaging amplification uses Bayesian networks or ensembling to estimate amplification from the precision of integrals over given phase-space volumes. Differential amplification uses hypothesis testing to quantify amplification without any resolution loss. Applied to state-of-the-art event generators, both methods indicate that amplification is possible in specific regions of phase space, but not yet across the entire distribution.

        Speaker: Sascha Cassandra Diefenbacher (Heidelberg University (DE))
      • 11:40
        Neural Control Variates for LO and NLO 20m

        We employ neural control variates to minimize the range of event weights and avoid negative weights for phase-space integration and event generation. A signed control variate, built from two normalizing flows, fulfills both tasks. Combined with neural importance sampling, it significantly reduces the computational cost of LO and NLO predictions. For the NLO case, our conditional neural control variate can be viewed as a trainable subtraction term, complementing the established physics subtraction schemes for enhanced sampling performance.

        Speaker: Sophia Vent (Institute for theoretical physics Heidelberg)
    • 11:00 12:00
      Parallel Talks: Reconstruction / Calibration
      • 11:00
        Amber: Self-Attention Muon Track Reconstruction in JUNO 20m

        Large-volume liquid scintillator detectors produce variable-size sets of PMT(photomultiplier tube) hits with rich timing, charge, and spatial correlations, making track reconstruction a natural application for attention-based architectures. This talk presents Amber (Attention Mechanism Based Event Reconstructor), a self-attention-based framework for cosmic-muon track reconstruction in the Jiangmen Underground Neutrino Observatory (JUNO)—a 20-kt liquid scintillator detector designed primarily to determine the neutrino mass ordering (NMO). Accurate and fast muon reconstruction is vital for JUNO to effectively reject cosmogenic 9Li/8He backgrounds, directly safeguarding the sensitivity of NMO determination and precision physics analyses.
        Amber represents each event using hit-level information from fired PMTs, where each hit is encoded by its first-hit time, collected charge, and PMT position, and a Transformer encoder learns global spatiotemporal correlations across the detector response to map the event-level representation directly to muon-track parameters. The framework provides two primary task-specific reconstruction models: Amber-S reconstructs single through-going muons as one straight track, parameterized by a reference point and a direction vector, while Amber-D extends the reconstruction to bundle-muon events using a double-track parameterization with a permutation-invariant loss. Additionally, Amber-OCR serves as an architectural variant of the backbone, specifically designed for high-multiplicity Central Detector PMT inputs by partitioning detector hits into local windows and combining local attention, convolutional feature extraction, and global attention for efficient detector-response modeling.
        Trained using a data-driven strategy, Amber directly regresses track parameters from detector hits and achieves an inference time below one second per event in production profiling, demonstrating the feasibility of deploying attention-based reconstruction in large-scale JUNO data processing.

        Speaker: Xiaoying Lu
      • 11:20
        Learning pairwise associations for event slicing in ProtoDUNE-SP 20m

        Event reconstruction in liquid argon time projection chambers (LArTPCs) requires separating detector activity from overlapping beam particles and cosmic rays into physically meaningful particle hierarchies. In the Pandora event reconstruction chain, event slicing groups particle-flow particles (PFPs) into candidate hierarchies using topological association criteria. We investigate whether this task can instead be formulated as a pairwise classification problem using simulated ProtoDUNE-SP data. Each event is represented as a graph in which reconstructed PFPs are nodes and edges encode the probability that two PFPs originate from the same true hierarchy. A graph neural network (GNN) is trained using reconstructed pair-level features and then evaluated based on Monte Carlo (MC) ground truth information. The GNN performance is compared with Pandora and machine learning (ML) baseline models. These findings motivate further investigation of deep learning (DL) adoption on event slicing in LArTPC data.

        Speaker: Christina Karagianni (University and INFN Ferrara (IT))
      • 11:40
        Optimizing Optimal Transport-based Pileup Mitigation 20m

        The current run of the Large Hadron Collider (LHC) yields on average 30-50 simultaneous pileup vertices per event, consisting of both charged and neutral showers. This number is expected to only increase at the High Luminosity LHC with predicted averages on the order of 140 pileup vertices. Pileup presents a salient problem that, if not checked, hinders the search for new physics and Standard Model precision measurements such as jet energy, jet substructure, missing momentum, and lepton isolation. While charged pileup is relatively simple to remove due to tracking information, neutral pileup remains a notable challenge for any LHC analysis. Current pileup mitigation techniques such as SoftKiller and PUPPI are all rule-based algorithms that require precise tuning, minimizing their overall flexibility. More recent work has focused on developing automated mitigation frameworks using machine learning. This study builds upon one such technique known as Training Optimal Transport using Attention Learning (TOTAL). The TOTAL methodology leverages a transformer architecture with an optimal-transport loss function to robustly learn an accurate description of pileup as a transport function by comparing matched samples with and without pileup interactions present, all without any need for assumptions of pileup nature. While TOTAL has already been shown to outperform conventional rule-based approaches, the extent of its potential performance remained unclear. In this work, we test the performance limits of TOTAL in two directions: 1) assessing the impacts of different configurations of input information and 2) testing new OT-based loss functions. Both directions have shown significant improvement over the original TOTAL methodology in terms of reconstructing relevant observables such as dijet mass for several BSM physics scenarios and a wide range of pileup conditions scaling up to 200 pileup vertices. These improvements further demonstrate the inherent value of using TOTAL as a new baseline for ML-based pileup mitigation.

        Speaker: Nathan Suri Jr. (Yale University (US))
    • 12:00 13:30
      Lunch Break 1h 30m
    • 13:30 15:30
      Parallel Talks: Generative Models
      Convener: Sascha Cassandra Diefenbacher (Heidelberg University (DE))
      • 13:30
        Efficient Event Generation for High-Multiplicity LHC Processes: An End-to-End GPU Workflow with Normalizing Flows 20m

        Producing very large unweighted event samples for high-multiplicity processes is limited by expensive matrix-element evaluations and low unweighting efficiencies. We present the first end-to-end GPU-resident event-generation workflow that integrates normalizing-flow proposals with the parton-level event generator Pepper. Helicity-conditioned coupling flows are trained using online updates supplemented by sample replay and deployed across all subprocesses of complete proton--proton collision processes with many final-state jets. In this workflow, a Python-based control layer and \Pepper exchange flow-generated phase-space points and the corresponding target-density evaluations directly in device memory. The control layer performs flow sampling, proposal-density evaluation, and unweighting, while \Pepper evaluates the matrix elements, PDFs, and phase-space factors defining the target density and writes the accepted events in standard formats. We compare subprocess-specific flows, with one flow per partonic subprocess, to grouped conditional flows that share parameters among subprocesses with related parton content. The workflow is benchmarked for $pp \to e^+e^- + 4j$, $pp \to e^+e^- + 5j$, $pp \to t \bar t + 4j$, $pp \to 4j$, and $pp \to 5j$ production. On four H100 GPUs, we generate (10^9) unweighted events for each benchmark process. Including the cost of flow training, the workflow achieves end-to-end speedups of up to two orders of magnitude over standalone \Pepper event generation and turns a multi-week task into a sub-day computation. It thereby makes billion-event production more practical and offers a pathway to alleviating the Monte Carlo statistics bottleneck in high-multiplicity collider physics.

        Speaker: Daohan Wang (HEPHY ÖAW)
      • 13:50
        Too good to go: Upcycling Phase-Space Points for Multijet Processes 20m

        Accurate sampling of multi-particle phase spaces is a major bottleneck in
        high-energy-physics simulations at the LHC, especially for processes with
        many final-state particles where matrix-element evaluations become
        prohibitively expensive. We introduce a stacked training strategy for
        phase-space point generators that cuts the training cost dramatically while
        delivering improved sampling performance. This method exploits the
        nested structure of a factorised phase-space parametrisation
        $\Phi_{N+1} = \Phi_{N} \times \Phi_{1}$
        property of the $(N+1)$-particle phase-space, so that a higher-multiplicity
        phase space can be constructed from a lower-multiplicity one, effectively
        transferring the knowledge of a low-dimensional sampler to a
        higher-dimensional one. We demonstrate the approach on
        Drell-Yan and $gg \to t\overline{t} + \text{gluons}$
        processes. Because the augmentation is used only to initialise the proposal
        distribution, exact event weighting preserves unbiased Monte Carlo estimates
        of physical observables.

        Speaker: Konrad Helms (Georg August University of Göttingen)
      • 14:10
        MadNIS at NLO 20m

        We combine fast amplitude surrogates with neural importance sampling to accelerate NLO calculations. For virtual corrections, a learned ratio to the Born matrix element with calibrated uncertainties guarantees reliable precision across phase space. For real emission, we stick to the standard FKS subtraction and train sector-conditioned surrogates of the regularized integrands away from divergences. MadNIS then uses multi-channel mappings and FKS sectors as conditions. We validate our approach for electron-positron scattering to three and four jets and find significant speed-ups and variance reduction in the integration.

        Speaker: Giovanni De Crescenzo
      • 14:30
        Sampling NNLO QCD phase space with normalizing flows 20m

        Next-to-next-to-leading-order QCD calculations are essential for precision collider physics, but their computational cost is often dominated by inefficient phase-space integration. In this talk, I will present the application of neural importance sampling to all contributions entering an NNLO QCD calculation of gluonic top-quark pair production within the STRIPPER subtraction framework. The approach compares discrete coupling-layer and continuous normalizing flows, with the samplers conditioned on sector and helicity information. A central improvement is the stratification of signed integrands into positive and negative components, followed by separate optimization of the corresponding sampling densities. The resulting models substantially reduce event-weight variances, improve unweighting efficiencies, and reproduce differential distributions consistently with conventional integration methods. When integrand evaluation dominates the runtime, the computational cost of reaching a fixed statistical precision is reduced by up to a factor of eight, demonstrating the potential of flow-based sampling for future high-precision collider calculations.

        Speaker: Timo Janßen (University of Göttingen)
      • 14:50
        Re-Ranking the CaloChallenge 2022 20m

        We present a new way of measuring the performance of generative models and use it to rerank all CaloChallenge submissions. This method reorders some submissions and provides a new perspective on how the generative models differ from the reference samples.

        Speaker: Agni Purani (Rutgers University)
      • 15:10
        FastSimu: Constraint-Preserving Generative Fast Simulation of Optical-Photon Response at Wavelength-Shifting-Fiber Entry 20m

        Optical-photon tracking is a major computational bottleneck in segmented plastic-scintillator detectors, where a single particle crossing can produce tens of thousands of photons. We present FastSimu, a conditional generative surrogate for the response of a 25-mm scintillator voxel at the first entry of photons into three orthogonal wavelength-shifting fibers. Conditioned on a multidimensional muon state vector, FastSimu generates per-fiber photon counts, timing summaries, and spatial distributions.

        The model is designed around the structure of the detector response: a heavy-tailed multivariate density model captures correlated scalar quantities, a diffusion model generates normalized spatial profiles, and multinomial sampling converts the generated totals and profiles into integer histograms with exact per-fiber count conservation. This combination provides physically valid outputs without relying on bin-wise independent predictions.

        FastSimu is evaluated on 200,000 Geant4 events. On held-out data, it closely reproduces the principal Geant4 distributions of photon counts, timing, and fiber-entry positions, with an average spatial discrepancy well below one histogram bin. In paired single-core benchmarks on the same workers, FastSimu achieves roughly a $4\times10^{3}$ improvement in amortized throughput at batch size 100.

        Because the voxel response can be evaluated independently and in parallel, FastSimu provides a scalable building block for hybrid simulation of large scintillator arrays. Downstream wavelength shifting, fiber transport, and SiPM response remain future extensions.

        Speaker: Mr Jiacheng Wu (Shanghai Jiao Tong University)
    • 13:30 15:30
      Parallel Talks: Reconstruction / Calibration
      • 13:30
        Learning to Reconstruct Muon Detector Showers from Raw Detector Readout 20m

        Long-lived particles decaying within the CMS muon system can deposit dense showers of hits in the endcap Cathode Strip Chambers (CSCs), called muon detector showers (MDS), whose hit multiplicity is the primary handle for their identification. Standard CSC reconstruction, built for isolated muon tracks, breaks down here: overlapping detector signals inflate and smear the hit count where the shower is densest. We present a geometry-aware cross-attention transformer that reconstructs shower hits directly from the raw CSC readout, trained on Geant4 truth hits by set prediction. It recovers the true hit multiplicity and cluster shape where classical reconstruction saturates, extending CSC hit reconstruction to dense showers that standard algorithms cannot handle.

        Speaker: Asu Guvenli (University of Hamburg (DE))
      • 13:50
        Probabilistic, Structure-Intrinsic Clustering with an Hierarchical Embedding (PSICHE) in Time and Space Applied In Particle Physics Jet Reconstruction 20m

        Jet reconstruction is an active and open area of particle physics research, with challenges and questions related to jet size, multiplicity, substructure, and experimental performance in the presence of pileup and noise. A new algorithm, Probabilistic, Structure-Intrinsic Clustering with an Hierarchical Embedding (PSICHE, arxiv:2608.xxxx), introduces a variety novel features to the domains of jet clustering and jet substructure, addressing several challenges and shortcomings of existing methods. These features include, but are not limited to: (i) dynamically-learned and variable jet sizes, (ii) unsupervised, multi-scale learning of emergent features, like jet substructure and multiplicity, (iii) clustering in space and time, and (iv) incorporation of domain-specific experimental uncertainties. All of these properties are achieved in a self-contained, probabilistic framework that is computationally tractable.

        This contribution will introduce and describe the PSICHE algorithm, with a range of examples at LHC energies for various physics phenomena (QCD jets, resolved/boosted top quarks, boosted Ws), clustering inputs (calorimeter cells, particle candidates), pileup scenarios (current and HL-LHC projected), and detector performance parameters.

        Speaker: Margaret Rose Lazarovits (The University of Kansas (US))
      • 14:10
        Transformer-based machine learning using low-level calorimeter signals for collimated photon identification 20m

        Electromagnetic calorimeters provide essential information for reconstructing and selecting both Standard Model (SM) and potential beyond the SM physics events at high-energy particle colliders. The fine-grained segmentation of modern calorimeters captures rich information about the internal structure of particle showers, much of which is discarded by conventional high-level reconstruction methods. In this talk, we show how machine learning applied directly to calorimeter cells in an ATLAS-like calorimeter can recover this information for a challenging benchmark task: distinguishing highly collimated diphoton signatures from light axion-like particle decays against isolated single-photon showers. We present a systematic comparison of six machine-learning architectures, from shower-shape-based approaches to direct cell-level methods. Cell-level learning delivers substantially better classification, with a Transformer achieving the best overall performance and an MLP Mixer offering a lightweight alternative suited to real-time, trigger-level use. We also show that the Transformer enables invariant mass regression directly from cells, sharpening the characterization of light resonances and providing a new handle against $\pi^0$ and $\eta$ fake photon backgrounds. Together, these results point to cell-level machine learning as a way to push calorimeter-based particle identification well beyond current techniques.

        Speaker: Gabriel Matos (Columbia University (US))
      • 14:30
        Transformer-based reconstruction of electromagnetic showers in CMS 20m

        The reconstruction of electrons and photons in the CMS Electromagnetic Calorimeter (ECAL) currently relies on a geometrical clustering algorithm called PFClustering. While it is efficient for isolated particles, it has a limited ability to resolve close-by showers and mitigate detector noise, which reduces the sensitivity of physics analyses and will worsen with detector ageing. We present ClusTEX, a single-step graph-attention-based transformer that reconstructs photon energy and position directly from calorimeter readout. Through a novel positional encoding scheme, ClusTEX remains sensitive to the relative graph structure of the events while simultaneously learning detector-position dependencies. We demonstrate that this approach eliminates the need for a dedicated classification step and remains robust to complex event topologies and non-responsive detector regions. Trained and evaluated on a toy ECAL simulation, ClusTEX improves upon both PFClustering and GNN-based approaches. Particularly for events with overlapping showers, ClusTEX significantly outperforms PFClustering in signal efficiency and per-photon energy resolution. The impact of this improvement is reflected in the reconstruction of boosted $\pi^0\rightarrow\gamma\gamma$ decays, where ClusTEX successfully resolves pions above 30 GeV that are inaccessible to PFClustering. Beyond boosted $\pi^0$ reconstruction, this approach provides a promising foundation for jet reconstruction and improved prompt-photon identification in $\gamma+\text{jet}$ analyses. We report on the ongoing integration of ClusTEX into the CMS Software (CMSSW) and present first results using the full detector simulation for the LHC Run 3.

        Speaker: Yuliia Maidannyk (IRFU-CEA, Université Paris-Saclay (FR))
      • 14:50
        Reconstruction of Tracker Material with Information Field Theory 20m

        Accurate knowledge of tracker material is essential for particle reconstruction and detector simulation in high energy physics. Photon conversions and secondary nuclear interactions provide localized probes of detector structures, but inferring a material map from their sparse and noisy vertex distributions is an ill-posed inverse problem. We present a Bayesian field-inference approach in which the spatial material distribution is modeled as a continuous latent field and reconstructed using Information Field Theory. A probabilistic forward model relates this field to the observed conversion and interaction vertices, naturally incorporating spatial correlations and enabling uncertainty quantification. The method is currently being developed and validated using configurable Geant4-based simulations, for which the inferred material maps can be compared directly with simulation truth.

        Speaker: Alexander Pesch Berrocal (Rheinisch Westfaelische Tech. Hoch. (DE))
      • 15:10
        Curvature-Adaptive Graph Neural Networks for Low-pT Charged-Particle Track Reconstruction at the HL-LHC 20m

        Charged-particle track reconstruction is a central challenge for the High-Luminosity Large Hadron Collider (HL-LHC), where high detector occupancy and the strongly curved trajectories of low-$p_T$ particles create severe combinatorial ambiguities. Although Graph Neural Networks (GNNs) have emerged as a promising approach for graph-based tracking, their performance is often limited by the quality of the underlying graph representation. We propose Manifold Helical-IN, a physics-informed framework that incorporates detector geometry and charged-particle trajectory information throughout the reconstruction pipeline. Candidate hit pairs are constructed by connecting only detector layers compatible with charged-particle propagation; pseudorapidity and azimuthal constraints reject geometrically inconsistent combinations, while the search region is adaptively enlarged according to the expected track curvature to retain low-$p_T$ trajectories without substantially increasing combinatorial background. The selected detector hits are then projected onto a helical manifold representation that follows the expected motion of charged particles in the solenoidal magnetic field, enabling the interaction network to distinguish genuine track segments from spurious connections more effectively. During message passing, information exchanged between neighboring hits is further weighted according to their geometric consistency with the underlying helical trajectory. Evaluated on the TrackML challenge dataset, Manifold Helical-IN achieves a reconstruction efficiency of 0.9675 in the benchmark region of $p_T > 500$ MeV, while in the challenging $p_T$ range of 380-460 MeV efficiency was 0.8496. The fake-track rate achieved was 0.0037, corresponding to a 13-fold improvement over the baseline Helical-IN architecture in the fake rate, while maintaining an inference latency of approximately 150 ms per event on a single NVIDIA A100 GPU. These results demonstrate that explicitly embedding detector geometry and charged-particle trajectory constraints into graph neural networks can substantially improve track purity while preserving computational efficiency, highlighting the potential of geometry-aware graph learning for scalable charged-particle track reconstruction at the HL-LHC.

        Speaker: Mr Syed Haider Ali (Institute of Physics, Faculty of Science and Technology, University of Debrecen, Egyetem tér 1, H-4032 Debrecen, Hungary)
    • 15:30 16:00
      Coffee Break 30m
    • 16:00 17:00
      Parallel Talks: Generative Models
      • 16:00
        Toward Complete ML-Based Simulation for CMS with FlashSim 20m

        FlashSim is an end-to-end machine-learning simulation in CMS that produces analysis-level events (NanoAOD) directly from generator-level input, at a small fraction of the cost of detailed simulation. Each reconstructed object is generated by its own model, continuous normalizing flows trained with flow matching, conditioned on generator-level information and the per-object models are combined at inference to build complete events, keeping the physical and identification variables correlated as in data. Recent developments make the output more complete and bring it closer to what analyses actually use. We show comparisons with fully simulated samples on analysis-level observables, indicating that FlashSim now covers more of what real analyses need.

        Speaker: Prabhat Solanki (Universita & INFN Pisa (IT))
      • 16:20
        To surrogate or to differentiate? A Delphes3 optimization study 20m

        Fast parametric detector simulations transform generator-level particles into reconstructed physics objects through a chain of modules controlled by smearing functions. A smearing function contains closed-form resolution and efficiency formulae with numeric coefficients which are traditionally tuned by hand against full simulation or data. We study gradient-based optimization of Delphes3, a multipurpose fast detector simulator for phenomenology studies, using a novel differentiable workflow that can optimize the full set of simulator parameters. We compare this approach to a surrogate network which combines simulation and reconstruction steps. The two workflows provide complementary views on the accuracy and interpretability of fast simulators for the LHC and future experiments.

        Speaker: Luigi Favaro (Universite Catholique de Louvain (UCL) (BE))
      • 16:40
        Parnassus for Detector Simulation with Legacy e⁺e⁻ Collider Data 20m

        Parnassus is a generative detector-simulation model that learns the mapping from generator-level particles to reconstructed particle-flow objects. We apply it to simulated hadronic Z-boson events from legacy e⁺e⁻ collider experiments. Using ALEPH simulation, we validate Parnassus across particle-, jet-, and event-level observables, including particle identification, reconstructed vertices, jet substructure, visible mass, and thrust. The model reproduces the full-simulation reference accurately, particularly for detector-sensitive and correlated observables. We further apply Parnassus to SLD by training a dedicated model on SLD simulation, with planned studies using DELPHI samples for additional cross-experiment comparison. This is the first study within Parnassus to transfer a generative detector-simulation model across different legacy e⁺e⁻ detectors.

        Speaker: Ya-Feng Lo (University of California Los Angeles (US))
    • 16:00 17:00
      Parallel Talks: Reconstruction / Calibration
      • 16:00
        Graph Neural Networks for Calorimeter-Based Cluster Calibration at the ATLAS Experiment 20m

        The ATLAS calorimeter system measures the energy of particles. These are large volumes segmented into many individual cells, each recording a small deposit of energy as a particle passes through. The cell-by-cell structure has a natural geometric and topological form well suited to graph neural networks (GNNs), a class of machine learning models designed to learn from data with irregular, connected structure rather than a fixed grid. In ATL-PHYS-PUB-2022-040, the ATLAS Collaboration showed that by using calorimeter information alone, GNNs can accurately perform calibration on topoclusters originating from both neutral and charged pions.

        This presentation extends that approach to topoclusters arising from more complex event topologies, including neutral hadrons and dijets. It further applies the technique to partially-subtracted clusters, which arise in ParticleFlow, a reconstruction strategy that combines calorimeter and tracking-detector information to improve energy measurement by removing the contribution of charged particles already measured elsewhere. The resulting calibration scheme demonstrates improved energy resolution and linearity relative to the baseline local hadronic calibration across the studied topologies.

        Speaker: Matthew Green (Adelaide University (AU))
      • 16:20
        Calibration of signatures from simulated electromagnetic showers in the CMS calorimeter with machine learning 20m

        Monte Carlo simulations are used extensively in high-energy particle physics analyses. However, imperfections in the configuration of detector simulation can lead to significant discrepancies between simulated events and collision data. Such mismodelling is often addressed using scale factors, which can be accompanied by large systematic uncertainties that compromise the sensitivity of measurements and searches. To mitigate potential biases and uncertainties arising from mismodelling, it is essential to calibrate simulations to data. We present two novel calibration methods based on machine-learning techniques: a reweighting approach, which employs a classifier to learn the ratio of probability density functions between simulation and data, and a normalising-flow approach, which learns a high-dimensional transformation to map simulation to data. Compared to traditional calibration methods, both approaches offer continuous, unbinned corrections across high-dimensional feature spaces, enabling improved global agreement with data. The techniques are demonstrated in the context of correcting the signatures of simulated electromagnetic showers in the CMS calorimeters, using proton-proton collision data collected during 2022 at $\sqrt{s}$ = 13.6 $TeV$, corresponding to an integrated luminosity of 26.7 $fb^{-1}$. The complementary strengths and limitations of the two methods are compared.

        Speaker: Caio Cesar Daumann (Rheinisch Westfaelische Tech. Hoch. (DE))
      • 16:40
        Latest machine learning techniques for reconstructing and calibrating hadronic objects in ATLAS 20m

        The precision and reach of physics analyses at the LHC is often tied to the performance of hadronic object reconstruction & calibration, with any incremental gains in understanding & reduced uncertainties being impactful on ATLAS results. Recent improvements from machine learning methods include the calibration & pileup tagging or calorimeter clusters, regressions of b-jet and boosted jet calibrations and their performance in data, and the ML-assisted reconstruction of missing transverse momentum.

        Speaker: Matthew Green (Adelaide University (AU))
    • 17:00 18:00
      Keynote 1h
    • 09:00 10:00
      Plenary Presentations
      • 09:00
        Overview Foundation Models 30m

        TBA

        Speakers: Anna Hallin (University of Hamburg), Anna Maria Cecilia Hallin (Universität Hamburg)
      • 09:30
        From Prompts to Publications: Agentic AI for Collider Analysis 30m

        Experimental high energy physics analysis is traditionally a multi-year effort dominated by repetitive implementation that demands little physics insight. LLM-based AI agents can now autonomously execute substantial portions of this pipeline, collapsing the implementation bottleneck to roughly ten hours of wall-clock time. This talk surveys the rapidly developing landscape of agentic AI in high energy physics, anchored by Just Furnish Context (JFC), a framework built on Claude Code that integrates autonomous analysis agents, literature-based knowledge retrieval, and a multi-agent review system mirroring the tiered review structure of a real collaboration. Given only a short physics prompt and a dataset, JFC plans and executes the full chain: event selection, background estimation, systematic uncertainty quantification, blinded statistical inference, and publication-grade documentation, with human oversight concentrated at a single unblinding gate. Demonstrations on ALEPH, DELPHI, and CMS open data include a CMS Run-1 $H\to\tau\tau$ measurement consistent with the published value and the first primary Lund jet plane density measured in $e+e-$ collisions, to our knowledge the first novel HEP measurement produced autonomously by an AI agent.

        I will situate these results among other recent agentic efforts across the community and confront the question these systems force: how do we know an agent's analysis is correct? I will present ongoing benchmarking work evaluating the robustness and physics fidelity of agentic analyses across models and tasks, and argue that as implementation cost collapses, the bottleneck shifts from coding capacity to physics ideas, review bandwidth, and trust.

        Speaker: Eric Anton Moreno (Massachusetts Inst. of Technology (US))
    • 10:00 10:30
      Coffee Break 30m
    • 10:30 12:30
      Parallel Talks: Fast Simulation
      • 10:30
        Simulating the CMS High-Granularity Calorimeter with Generative AI 20m

        In the upcoming High Luminosity LHC era, detector simulation will face computing resource constraints; at the same time CMS will be upgraded with the new High Granularity Calorimeter (HGCal), which is more intensive to simulate. This computing challenge motivates the use of generative machine learning models as surrogates to replace full physics-based simulation of particle showers in the HGCal. A large dataset of calorimeter showers in the CMS HGCal, simulated with GEANT4, has been prepared and will be used to train and assess the performance of multiple state-of-the-art generative AI models. Applying generative AI to the simulation of HGCal showers requires significantly higher granularity and more irregular geometry as compared to previous AI-based calorimeter simulation studies. We will discuss the various methods employed by the AI models to overcome these challenges. The quality of the showers produced by the various AI models will be assessed and compared across multiple performance metrics.

        Speaker: Sitian Qian (Northwestern University and Fermilab)
      • 10:50
        Fast Calorimeter Simulation in ATLAS with Modern Generative Models 20m

        Simulating electromagnetic and hadronic showers accurately is among the most computationally demanding parts of the ATLAS detector simulation. To cut CPU consumption for Run 3, the collaboration deployed AtlFast3, a fast simulation tool that pairs traditional histogram-based parameterisations with GAN-based calorimeter models. For the upcoming  Run 4 of the LHC, work began on optimising the voxelisation scheme used to train these models, which organises energy deposits into small volumetric bins. This revised voxelisation yields a more efficient description of showers and a marked gain in physics performance. Even with these improvements, GANs continue to exhibit their familiar shortcomings in stability and accuracy. For this reason, ATLAS is exploring more recent generative techniques, including diffusion models, transformers, and continuous normalizing flows as suggested by the recent CaloChallenge. Early findings indicate that these models capture fine-grained shower characteristics more dependably also on real world data, and they are being assessed as candidates to replace portions of the parameterisation currently used in AtlFast3. This contribution reviews the state of machine-learning-based calorimeter simulation in ATLAS, reports results from the modern generative models under investigation, and examines the principal obstacles on the route to a more ML-driven fast simulation for Run 4 and beyond.

        Speaker: Florian Ernst (Heidelberg University (DE))
      • 11:10
        AllShowers: One model for all calorimeter showers 20m

        To address the high computational demand of detector simulations in high-energy physics, various generative surrogate models have been developed. Traditionally, one generative model per incident particle type is trained, requiring separate trainings and model weights. This increases training and human effort, as well as memory footprint during inference, since multiple sets of weights must be loaded. To address this, we put forward AllShowers, a single generative surrogate capable of generating showers from 12 different particle types, supporting a wide range of incident angles and energies, without the need for retraining or fine-tuning. The model receives the incident particle type and kinematics as conditional input. We demonstrate high-fidelity generation of electrons, photons, and charged and neutral hadrons in the highly granular electromagnetic and hadronic calorimeters of the International Large Detector (ILD). In addition to unifying the generation, AllShowers surpasses the fidelity of previous state-of-the-art single-particle-type models for hadronic showers in highly granular calorimeters.

        Speaker: Thorsten Buss (RWTH Aachen)
      • 11:30
        Fast Point Cloud Generation via Latent Diffusion 20m

        Fast, reliable surrogates for detector simulation are essential to support the physics program of the high-luminosity LHC (HL-LHC) and future collider experiments. We present a generative model for pion showers in the ECal and HCal of the International Large Detector (ILD), representing each shower as a high-granularity point cloud where every point encodes position, energy, and — for the first time for hadronic showers — timing information. The model avoids the cost of direct point-cloud generation by encoding every shower into a latent representation: a diffusion process produces compact latent codes, which a VAE decoder then unfolds into full point clouds only in its final layers. Therefore, this approach is significantly cheaper than direct point-cloud generation, while still producing showers with realistic spatial and temporal structure. The latent-space formulation offers a general strategy for scaling point-cloud generative models to the high-multiplicity showers produced by highly granular calorimeters, while also incorporating timing information.

        Speaker: Martina Mozzanica (University of Hamburg)
      • 11:50
        step2point: Optimising Point Cloud Representations for Fast Calorimeter Simulation 20m

        The substantial computational burden associated with the use of traditional Monte Carlo simulation for the calorimeter systems of high energy physics experiments has driven the development of numerous deep generative models for fast calorimeter simulation. Recently, several models have been proposed which move away from the common image-like representation of a shower using a regular grid to a more flexible point cloud representation of energy deposits. However, directly using the large number of energy deposits produced by detailed Geant4 simulations is computationally prohibitive. For this reason, previous work has used handcrafted, detector-dependent procedures to create point clouds with a sufficiently reduced number of points, while still preserving key physics observables. To address this challenge, we present the step2point library, a lightweight and configurable tool for preprocessing electromagnetic and hadronic showers into compact point-cloud representations, while preserving physically relevant shower characteristics.

        In this contribution, we demonstrate the functionality of the step2point library for detectors including the Open Data Detector (ODD), the International Large Detector (ILD), and the CLIC-like Detector (CLD), achieving significantly reduced data volumes while maintaining accurate simulation-level physics observables. We highlight the advantages of the step2point workflow with the example of the CaloClouds3 model trained on electromagnetic showers. Finally, we present initial studies aimed at producing point clouds for hadronic showers.

        Speaker: Peter McKeown (CERN)
    • 10:30 12:30
      Parallel Talks: Foundation Models
      • 10:30
        LI-JEPA: Self-supervised pretraining of Lorentz Invariant Models 20m

        Deep learning classifiers for high-energy physics have recently advanced along two largely separate tracks: self-supervised pretraining on large unlabelled datasets, and architectures that respect Lorentz symmetry by construction. How these directions combine has not been shown.

        To address this, we integrate Lorentz-invariant models into the LeJEPA pretraining framework, in which the network is trained to produce a latent representation that is invariant under predefined augmentations. This naturally aligns with the philosophy of Lorentz-invariant models, where invariance under augmentations consisting purely of a Lorentz transformation is trivially enforced by the network.

        We evaluate on JetClass under matched pretraining data, augmentations, and optimization, varying only the backbone. Under this comparison the Lorentz-invariant model achieves substantially superior linear separability and downstream tagging accuracy relative to a non-invariant scalar transformer, showing that exact physical symmetries and self-supervised pretraining are complementary rather than competing formulations.

        Speaker: Andreas Hermansen (Universite de Geneve (CH))
      • 10:50
        To Mask, Predict, or Classify? Comparative Analysis of Pre-Training Strategies for Jet Foundation Models 20m

        In recent years, several pre-training strategies have been proposed for foundation models in jet physics. These approaches range from self-supervised generative tasks like next-token prediction (NTP) and masked particle modeling (MPM), to standard supervised classification. Inputs to foundation models are often tokenized, but this leads to a loss of information. Recent work has shown that using a hybrid setup with continuous feature inputs can prevent this loss of information. However, finding the best way to combine these different pre-training objectives remains an open question. In this study, we investigate whether combining these diverse pre-training tasks translates to improved model efficacy on downstream applications. We systematically evaluate a variety of pre-training setups, including supervised classification alone, classification combined with MPM, classification with NTP, and a joint strategy utilizing all three objectives. We will highlight the general trade-offs between discriminative performance and generative capabilities to help guide the future development of robust collider physics foundation models.

        Speaker: Anuranjan Sarkar (DESY)
      • 11:10
        Performance of Tabular Foundation Models for analysis of 4-vectors LHC final states with the FAIR Universe HiggsML Dataset 20m

        In collider-based particle physics experiments, proton collision events are commonly represented as tabular datasets for specific final states, features being four-vectors of final state particles or higher-level variables (invariant masses).
        Inspired by the success of foundation models in language and vision, recent developments have introduced tabular foundation models such as TabPFN and TabICL.
        We explore the use of tabular foundation models in collider physics analyses, and we show that they clearly outperform baselines like XGboost in data-limited regimes, for 50K training events or less. PFNs are benchmarked against boosted decision trees and neural networks across classification, density ratio evaluation (including calibration), and generation using the FAIR Universe dataset. We discuss performance, data efficiency, and computational trade-offs, and assess the potential role of pretrained tabular models in HEP analysis workflows.

        Speaker: David Rousseau (IJCLab-Orsay)
      • 11:30
        What can representation geometry reveal about collider foundation models? Towards Scientific AI as a New Microscope 20m

        Foundation models have demonstrated remarkable performance in collider physics, yet little is understood about the geometric principles underlying their neural representations. We propose representation geometry as a new direction for Scientific AI, where machine learning models serve not only as powerful predictors but also as “microscopes” for revealing the intrinsic structure of the underlying physical data. In this work, we study the OmniJet family of collider foundation models pre-trained with three different self-supervised objectives on the JetClass benchmark dataset. Using the Gromov-Wasserstein distance, we compare the metric geometry of representation spaces across network blocks, training steps, and pre-training schemes. Our preliminary results uncover highly structured geometric evolution both across network depth and throughout training, while suggesting that learned representations exhibit both common structural features and systematic variations induced by pre-training objectives. These studies establish a framework for investigating whether and to what extent collider foundation models converge toward common geometric structures, laying the groundwork for future empirical tests of the Platonic Representation Hypothesis in a scientific domain whose underlying reality is governed by fundamental physical laws.

        Speaker: Tianji Cai
      • 11:50
        Scaling tokenized models for full-event generation 20m

        Full simulation and reconstruction are projected to become major bottlenecks for computation at the High-Luminosity LHC, motivating the need for fast, ML-based surrogates. At the same time, LLMs have driven fast progress in generative discrete modeling: autoregressive transformers trained on tokenized data now represent the state of the art across a range of generative tasks. We extend the discrete modeling paradigm by introducing a general-purpose particle-level generative model trained on tokenized full-event data. We demonstrate the ability of this model family to perform both conditional generation from detector-stable particles and unconditional generation; study its scaling behavior across dataset and model sizes, and show that the learned representations transfer to downstream tasks.

        Speaker: Dan Godi (Weizmann Institute of Science)
      • 12:10
        Top quark spin and CP studies with L-GATr 20m

        The top quark is well suited for precision tests of the Standard Model and also provides a unique window into physics beyond the SM, especially via probing of its spin state. Its spin information is encoded in its decay products and is needed for the construction of genuine CP-odd observables, making an accurate polarization reconstruction essential. In this work, we employ L-GATr to study the reconstruction of hadronic top quark polarization and examine CP violation in ditop production. By tagging down-type quarks, we achieve a hadronic spin analyzing power of 0.72. Furthermore, we demonstrate that L-GATr can be used to separate the CP-even and CP-odd information in a CP-violating process and apply it to the ditop production process.

        Speaker: Marco Menen (TU Dortmund)
    • 12:30 14:00
      Lunch Break 1h 30m
    • 14:00 16:00
      Parallel Talks: Fast Simulation / Explainable AI
      • 14:00
        SPADE: Split-and-Delay Embeddings for Autoregressive High-Granularity Calorimeter Simulation 20m

        We introduce SPADE (SPlit And Delay Embeddings), which embeds each feature of a token independently and staggers the resulting streams along the sequence with progressively increasing delays. The vocabulary then scales additively rather than multiplicatively, while intra-token correlations are recovered by the ordinary causal self-attention mechanism: each feature is predicted at its own sequence position, conditioned on the features already emitted for the same object. No auxiliary decoder or quantization stage is required.

        We demonstrate SPADE on point-cloud shower generation in the highly granular ILD electromagnetic calorimeter. SPADE is competitive with state-of-the-art flow-matching on photon showers and substantially outperforms VQ-VAE-based predecessors. Against a joint-vocabulary baseline at the finest granularity, SPADE uses 74× fewer parameters and converges 6.9× faster in GPU hours, while better reproducing observables sensitive to energy–position correlations.

        By eliminating the need for quantized codebooks, SPADE enables direct, LLM-style autoregressive pretraining on multi-feature sensor data across fundamental physics.

        Paper: https://arxiv.org/abs/2606.11304

        Speaker: Henning Rose (Uni Hamburg)
      • 14:20
        Improving neural generative models with density estimation and Monte-Carlo resampling 20m

        Detector simulation is among the most resource-intensive components of modern collider experiments, currently consuming roughly half of the LHC computing budget and more still in the High-Luminosity phase. Normalizing flows are attractive surrogate models for fast simulation: they sit on the Pareto frontier between inference speed and accuracy compared with competing generators such as VAEs and GANs, and support both generation and tractable density estimation. Nevertheless, closing the measurable fidelity gap—especially at high detector granularity—is necessary for flows to be truly competitive. In this study, we improve the generation quality of normalizing flows without modifying their architecture.

        First, we show that a classifier trained to distinguish reference from synthetic showers provides an effective estimator of the density ratio between the two distributions. Using this ratio with acceptance–rejection and Markov chain Monte Carlo (MCMC) sampling, we achieve better agreement with the reference distribution across a range of physically motivated high-level observables. This generic technique can be used in principle to improve any generative model.

        Second, exploiting the bijective nature of flows, we perform the same resampling directly in latent space, where it is much cheaper than in data space. A classifier-based density-ratio estimate between the latent and base distributions is used to resample the base distribution before the flow's backward transformation, achieving comparable fidelity gains with minimal additional time and compute.

        Speaker: Dr Minh Tuan Pham (IJCLab, Université Paris-Saclay, CNRS/IN2P3)
      • 14:40
        Information Geometry of Learned Latent Representations 20m

        Many classification and reconstruction tasks in physics rely on learned latent representations of the data. When networks are trained with a notion of locality, they encode task-specific similarity as closeness in the latent space. Differential geometry, particularly information geometry, is a powerful tool to uncover the learned information in these latent representations and thereby retrace the decision-making of the network. 
With a jet tagging task in mind, we construct an information geometry on the learned latent space of a variational autoencoder with a classifier head. We show how the classifier likelihood links high-level physics observables to the learned latent geometry via the Fisher information metric. We supplement the Fisher information by advanced differential geometric concepts such as curvature and geodesic paths and use them to assess the importance of different features for the classification decision. We realize the importance of the likelihood's skewness contribution as a geometric nonmetricity tensor and use it to construct coordinate-invariant measures of class separation.

        Speaker: Rebecca Maria Wolfgramm-Kuntz (Astronomisches Rechen-Institut, Zentrum für Astronomie, Universität Heidelberg)
      • 15:00
        Information Geometry of Latent Representations for Interpretable Jet Tagging 20m

        Latent representations are an important theme in modern machine learning. While information geometry provides a framework to analyze their structure, it also offers new insight into the physics learned by jet-tagging networks. We apply these methods to binary quark-gluon classification and three-fold fat-jet tagging and relate the learned latent representation to characteristic features of QCD radiation and heavy-particle decays. We show how information-geometric observables identify the relevant physics encoded by the network and connect the classifier’s decisions to established jet observables.

        Speaker: Benedikt Schosser
      • 15:20
        Where Does the Particle Transformer Look Inside a Jet? A Jacobian-Lens Interpretation 20m

        Particle Transformers have achieved state-of-the-art performance in jet tagging, but the physical information underlying their decisions remains difficult to characterize. This limits our ability to assess their robustness, diagnose sensitivities to Monte Carlo mismodeling, and design effective pretraining strategies. We introduce a Jacobian-lens framework that resolves the response of a Particle Transformer into contributions associated with individual jet constituents and input features. The resulting constituent- and feature-level profiles provide a local picture of where and how discriminating information is encoded throughout the network. By comparing models trained or pretrained on different datasets, we uncover nontrivial changes in the particles, kinematic regions, and physical features emphasized by the model. The Jacobian lens therefore provides a systematic diagnostic of learned representations in jet taggers, offering a route toward more interpretable, simulation-robust, and physics-informed transformer models.

        Speaker: Sitian Qian (Northwestern University and Fermilab)
      • 15:40
        Do Jet Taggers Learn Physics? Probing and Steering Latent Directions in a Transformer Tagger 20m

        A growing body of work on large language models has focused on steering vectors, which are linear directions within a model’s latent space that correspond to interpretable language concepts the model has learned. However whether such interpretable directions exist in models trained on physics data, that correspond to physics concepts, remains underexplored. We investigate this in the context of jet tagging, on a transformer architecture with no physics-motivated inductive biases trained on the JetClass dataset. We find that such steering vectors do seem to exist and we show specific directions that correspond to physically meaningful concepts like jet mass, constituent particle multiplicity, and jet momentum. We provide both correlational evidence through linear probing and causal evidence through steering vector interventions, demonstrating that these directions influence model predictions. For example, applying the constituent particle multiplicity direction as a steering vector to a batch of Hcc jets shifts the model's prediction toward Hgg. We also touch on broader potential applications and characteristics of steering vectors within jet tagging models, such as whether steering vectors can be used to calibrate synthetic data to better mimic observed data and whether these linear directions differ as you add inductive biases to the model.

        Speaker: Daniel Tiourine
    • 14:00 16:00
      Parallel Talks: Foundation Models / Agentic Workflows
      • 14:00
        Measurement of Spin and Quantum Correlations in $Z \to \tau^+\tau^-$ Decays at LEP with DELPHI 20m

        Precision studies of $\tau^+\tau^-$ production at the Z pole provide a clean environment for investigating electroweak spin correlations and quantum information observables. Using archived LEP-1 data collected by the DELPHI experiment, the process $e^+e^- \to Z \to \tau^+\tau^-$ is well measured, but the presence of multiple neutrinos in $\tau$ decays limits reconstruction of the $\tau$-pair rest frame. This constrains the precision of spin-dependent measurements.

        We present a machine-learning reconstruction method based on an event-level foundation model that predicts neutrino momenta using detector-level observables and kinematic constraints. The reconstructed $\tau^+\tau^-$ rest frame improves the resolution of spin-sensitive observables and enables a measurement of the $Z \to \tau^+\tau^-$ spin density matrix. These results provide a basis for quantum correlation studies using archived $e^+e^-$ collider data and demonstrate the applicability of machine-learning-based multi-neutrino reconstruction in precision electroweak analyses.

        Speaker: Mr Chen-Hua Hsu
      • 14:20
        Preference-Optimized Generative Models for Underconstrained Inference in Particle Physics 20m

        Precise reconstruction of event kinematics in the presence of invisible particles constitutes a fundamental underconstrained inference problem in particle physics. In processes such as dileptonic $t\bar{t}$ production, multiple undetected neutrinos lead to a multimodal solution space, where several kinematically consistent configurations can explain a single observed event.

        We present a preference-optimized generative framework for underconstrained inference, built on the event-level foundation model EveNet. Our approach augments a diffusion-based generative model with Direct Group Preference Optimization (DGPO), a post-training strategy that shifts the learned distribution toward physically preferred solutions while preserving its multimodal structure.

        Evaluated on the $t\bar{t}$ dilepton channel, the preference-optimized model improves reconstruction fidelity and reduces unfolded uncertainties compared to both the baseline EveNet diffusion model and $\nu^2$-Flows across key observables. These results establish preference optimization as an effective paradigm for generative inference in underconstrained systems, with direct implications for precision measurements at the LHC.

        Speaker: Mr Ting-Hsiang Hsu
      • 14:40
        Searching For Anomalies with Foundation Models 20m

        Foundation models have the potential to expand the discovery reach for new physics searches. In this paper, a full physics analysis using CMS Open Data is carried out to search for new physics in dijet events using the OmniLearned foundation model. We find that the background estimation describes the data well in validation regions, but is unable to accurately model the signal region where a significant excess is observed.

        Speaker: Vinicius Massami Mikuni (Lawrence Berkeley National Lab. (US))
      • 15:00
        Foundation Models for Jet Anomaly Detection: Optimization and Generalization 20m

        Foundation models have recently emerged as a promising approach for learning transferable representations from low-level particle physics data. In this work, we investigate the application of the OmniJet-α foundation model, to jet anomaly detection, focusing on how pretraining strategies and downstream optimization affect performance on the LHC Olympics (LHCO) benchmark. We systematically compare supervised and self-supervised pretraining objectives and explore transfer-learning strategies, including batch composition and model-selection criteria for downstream anomaly detection. Together, these studies identify effective approaches for transferring foundation models to jet anomaly detection. Beyond LHCO, we investigate the generalization of the optimized models using benchmark signal samples from the CMS anomaly detection analysis, establishing a benchmark for future studies of foundation models in fully agnostic collider searches. These findings highlight the potential of foundation models to enhance fully agnostic anomaly detection in collider physics.

        Speaker: Soumya Shaw (CISPA, Saarland University)
      • 15:20
        Agentic AI for Automated Foundation Model Analysis in HEP 20m

        Agentic AI frameworks offer a promising path toward automated HEP analyses, but current approaches rely on iterative prompting and accumulated context, leading to limitations in reproducibility and generality. We try to address these limitations by designing a structured agentic framework that drives a physics foundation model (FM) as its engine, replacing ad-hoc prompt engineering with systematic workflow steps.

        We demonstrate the framework using EveNet as a case study, on two benchmarks: a search for exotic Higgs boson decays (H→aa→4b) and a measurement of quantum correlations in dileptonic tt̄ production. Given a natural-language physics goal, the agent selects the appropriate task-specific components of the FM, confirms with the physicist, converts input data, monitors fine-tuning, and computes the target observable. We will present the results of both benchmarks and discuss the remaining challenges toward fully reproducible agentic HEP analyses

        Speaker: Yue Xu (University of Washington (US))
      • 15:40
        ColliderBench: Benchmarking AI Agents with Particle Physics Analysis Reproduction 20m

        Autonomous language-model agents are increasingly evaluated on long-horizon tool-use tasks, but existing benchmarks rarely capture the complexity and nuance of real scientific work. To address this gap, we introduce ColliderBench, a benchmark for evaluating whether LLM agents can reproduce parts of experimental analyses from the Large Hadron Collider (LHC) using only public papers and open scientific software. Such analyses are often difficult to reproduce because the public toolchain only approximates the software used internally by the experimental collaborations, while the published papers inevitably omit implementation details needed for a faithful reconstruction. Agents must therefore rely on physical reasoning, domain knowledge, and trial-and-error to fill these gaps.

        Speaker: Sofia Palacios Schweitzer (Rutgers University)
    • 16:00 16:30
      Coffee Break 30m
    • 16:30 18:10
      Parallel Talks: Agentic Workflows
      • 16:30
        MadAgents 20m

        We present MadAgents, an effective and communicative set of agents for working with MadGraph. Agentic installation, learning-by-doing training, user support, and autonomous simulation campaigns provide easy access to state-of-the-art simulations and accelerate LHC research. We show how MadAgents interact with inexperienced and advanced users, support a range of simulation tasks, and analyze the results. We also discuss the updated MadAgents implementation, with significantly improved answer reliability and the ability to improve through use.

        Speaker: Daniel Schiller (Institute for Theoretical Physics, Heidelberg University)
      • 16:50
        Agentic AI applications to particle theory and collider phenomenology 20m

        I will review the first generation of agentic AI tools for high-energy theory and simulations, including HEPTAPOD for orchestrating collider phenomenology workflows and Diagrammatica for autonomous symbolic Feynman-diagram calculations. I will also describe the ongoing community efforts for creating an ecosystem-level infrastructure for deployability, reproducibility, community uptake, and stewardship of such tools in the future.

        Speaker: Konstantin Matchev (University of Alabama (US))
      • 17:10
        Ariadne: Agentic Automation of Collider-Physics Workflows 20m

        LLM agents increasingly assist with coding, but still struggle with long, tool-heavy, multi-step tasks common in high-energy physics. We present Ariadne, an autonomous multi-role LLM agent that, given a natural-language task, is designed to plan, divide, and execute this workflow step by step.
        Ariadne's orchestration is based on a LangGraph state that organizes the workflow context into structured, persistent records and does not rely on an unstructured conversation history. Coupled with a context manager, every LLM prompt is curated for the relevant node and task, using the workflow state and the current step's requirements. Targeted retrieval from supplied papers, arXiv articles, and documentation provides the context manager with relevant source material without loading full sources into every prompt. Artifact and execution records preserve provenance and support reproducibility, and node-specific prompts limit context degradation across the workflow. Ariadne validates results during the workflow and can debug errors through bounded repair and replanning loops. It integrates tools including MadGraph5, Pythia 8, Delphes, ROOT/uproot, and Prospino, and has native support for execution on local, Slurm, or HTCondor environments.
        The framework is compatible with commercial or self-hosted LLM backends, enabling cost-conscious deployment. We discuss agent design and collider-analysis tests, emphasizing context engineering and validation to support reliable LLM agents for scientific computing.

        Speaker: Aman Upadhyay (Rutgers University)
      • 17:30
        Agentic Re-Casting using Agentic Re-Simulations 20m

        Analysis re-casting at the LHC is highly standardized and nevertheless requires resources, time, and expert physics input. Building on the newly developed MadAgents.v3 framework, we demonstrate how a global SMEFT analysis can be updated and improved through an agentic workflow with a physicist in the loop. Although demonstrated within the SFitter framework, the underlying technical aspects of the agentic recasting we present can be readily extended to other analysis frameworks.

        Speaker: Nikita Schmal (Universität Heidelberg)
      • 17:50
        Optimizing local LLM agents for HEP 20m

        LLM agents are beginning to take on real high-energy-physics workflows — tasks that demand orchestrating full toolchains of event generators, detector simulation, and analysis code rather than simply producing text. Making such agents reliable, auditable, and cheap enough to run on local hardware raises a distinct set of questions from those studied in general-purpose agent benchmarks.

        This talk examines what currently limits these agents, how their failure modes shift with model capability, and where the computational cost of an agent loop actually accumulates — often not where conventional inference optimizations target it. On the quality side, we discuss scaffold design and fine-tuning on collider-specific tasks, using agent trajectories and validated simulation artifacts as training data. On the speed side, we introduce speculative decoding, a technique that lets a model emit highly repetitive text much faster, and show why the templated structure of generator and detector configuration files makes it especially well suited to physics workflows.

        We discuss these questions concretely in the context of analysis-recasting benchmark tasks and locally served, in-house-optimized models.

        Speaker: Darius Faroughy (Rutgers University)
    • 16:30 18:10
      Parallel Talks: Explainable AI
      • 16:30
        Autoencoders for Symbolic Distillation of Black-Boxes 20m

        Deep learning (DL) approaches to high-energy jet tagging achieve state-of-the-art performance over classical methods, but lack human interpretability. We propose a framework for constructing post-hoc interpretable symbolic surrogates of state-of-the-art DL jet taggers trained on particle cloud input representations. A Deep Sets variational autoencoder first learns fixed-size tabular representations of the input data. These latent variables then serve as inputs to a black-box symbolic distillation step that models a DL model's decision boundary. Finally, a separate latent-space symbolic distillation step then interprets the latent space itself using known jet observables. We evaluate the framework on three benchmark jet tagging tasks found in the JetClass dataset across three state-of-the-art DL architectures. We find that surrogates tend to trade classification performance for interpretability in comparison with DL models. Additionally, we identify several functional similarities between surrogates that separate jets via N-subjettiness variables and terms representing energy distributions. The results of our framework suggest that it can be used as a general purpose approach to correlating particle cloud jet tagger decision boundaries to physically meaningful jet substructure properties.

        Speaker: Alan Gu
      • 16:50
        Jet Taggers Rediscover the Energy Correlators: A Causal and Explainable-AI Dissection 20m

        We open the black box of machine-learning jet taggers, asking how they compute their decisions and whether they rediscover QCD. Applying the causal mechanistic-interpretability toolkit (ablation, path patching, logit-lens, probing) to a Particle Transformer top-tagger, we isolate a sparse six-head source -> relay -> readout circuit that recovers 97.3% of full-model AUC, encodes the energy-correlator basis over N-subjettiness, and implicitly factorizes top tagging into two-prong W→qq̄ identification. A complementary physics-informed explainability study on the Lund Jet plane, three explainers across LundNet, ParticleNet, and ParT, over 1-/2-/3-prong tagging and p_T bins shows explainer importance tracking the same substructure observables (τ₂₁, τ₃₂, C₂, C₃) across architectures, confirming that taggers learn genuine energy-correlator physics.

        Speaker: Sanmay Ganguly (Indian Institute of Technology Kanpur (IN))
      • 17:10
        A Unified Geometric Framework for Understanding Collider Event Manifolds 20m

        As particle collider experiments produce increasingly large and complex datasets, a fundamental question arises: how should we quantify the similarity between two events? A variety of physically motivated metrics have been developed—from optimal transport to phase-space distances—yet the geometric structures induced by these metrics remain largely unexplored. We introduce the Multi-Reference Relative Representation (M3R), a framework that embeds diverse distances into a common coordinate space, thus enabling a systematic study of collider event manifolds. Using M3R, we investigate the geometry of single-event manifolds and the decision boundaries separating distinct physical processes. Further, multiple metrics are combined into unified event representations that capture complementary information. Our M3R framework provides an interpretable geometric language for organizing metric spaces, opening the possibility of a unified study of collider event geometry across analytical and data-driven approaches.

        Speaker: Hancheng Li (Rutgers University)
      • 17:30
        Interpreting Parton Distributions with Shapley Values 20m

        We show that Shapley values can be used to trace how individual parton distributions (PDFs) shape the theory predictions for high-energy observables computed from them.
        This provides a tool for assessing the impact of data on PDFs when determining them, and the impact of PDF uncertainties when using the PDFs to compute collider observables.
        The Shapley value is computed by treating the regression of PDFs from data as a cooperative game.
        The PDFs are the players, and the reward is the likelihood ($\chi^2$) that characterizes the agreement between data and the predictions obtained from a given PDF, with theory and methodology held fixed. The method is agnostic to the way PDFs have been determined in the first place: for PDFs determined with a black-box AI model it may be used in order to explain the behavior of the model, and for PDFs determined using a fixed parametrization it may be used in order to expose the features and potential limitations of the parametrization.
        We find that the method recovers known expectations about which data constrain which PDFs in a global fit, while placing them on a more quantitative footing.
        We demonstrate its effectiveness in two ways. We uncover an unexpected loss of sensitivity of the gluon PDF at intermediate $x$, with potential implications for BSM searches and the gluon fusion Higgs cross section. We also show that the method can be used to improve the hyperparameter optimization procedure currently used by the NNPDF collaboration.

        Speaker: Raphael Bonnet Guerrini (Computer Science Dep. University of Milan)
      • 17:50
        Data driven hadronization with HOMER 20m

        Due to the non-perturbative nature of hadronization, its simulation relies on a fragmentation function of a fixed parametric form. We present HOMER, a data-driven alternative based on neural networks that extracts the Lund string fragmentation function directly from data. HOMER addresses the information gap between the latent and observable phase spaces through an iterative reweighting procedure. We explore how far this approach can be pushed by increasing the complexity of the string configurations, which widens the information gap, and by departing from the assumption of the reference simulation.

        Speaker: Susie Kim (ITP, Heidelberg University)
    • 09:00 10:00
      Plenary Presentations
      • 09:00
        Anomaly Detection: From Methods to Applications 30m

        Traditional new-physics searches face a trade-off between coverage and sensitivity. Anomaly detection offers a way to extend their reach by searching for unusual signatures without assuming a specific signal model. Over the past decade, the field has developed a broad range of anomaly-detection paradigms and methods. More recently, the focus has begun to shift from developing new methods towards improving their sensitivity and coverage and, importantly, understanding what is required to make them useful in experiments, including questions of robustness, statistical interpretation, and integration into existing search strategies. In this talk, I will review key developments in anomaly detection and discuss the challenges and opportunities ahead as the field learns from past physics applications to improve future ones. I will also discuss the new anomaly detection task force in the BSM Working Group, which aims to consolidate ongoing efforts and help address remaining obstacles to the widespread use of anomaly detection in experiments.

        Speaker: Marie Hein (RWTH Aachen University)
      • 09:30
        Fast AI / Trigger Overview 30m

        TBA

        Speaker: Sioni Paris Summers (CERN)
    • 10:00 10:30
      Coffee Break 30m
    • 10:30 12:30
      Parallel Talks: Anomaly Detection
      Convener: Barry Dillon (b.dillon@ulster.ac.uk)
      • 10:30
        Getting More Power from Anomaly Detectors for Free with Weights 20m

        Weakly-supervised anomaly searches, such as CWOLA and CATHODE, typically train a classifier to separate a signal region from a sideband-estimate background and then cut on the classifier output to create a signal-enriched dataset. We show that there is a more efficient use of the output: the same classifier output can be used instead as a per-event weight, This results in a more powerful test with no cuts and significantly fewer discarded events. We formulate the anomaly detection problem as a hypothesis test over a scaled Poisson likelihood, and we show how to extract a profile-likelihood test statistic that is consistent with Wilks' Theorem under the null hypothesis, allowing for the precise extraction of meaningful $p$-values. We also provide theoretical estimates for the power of the weakly supervised anomaly detection tests, both traditional and weighted, placing the study of the statistical properties of anomaly detection on firmer theoretical ground. This weights method can be easily applied to any existing weakly supervised anomaly detection search for free,  with no additional training or calibration required.

        Speaker: Rikab Gambhir (University of Cincinnati)
      • 10:50
        Improving direct background estimation for resonant anomaly detection 20m

        Resonant anomaly detection is a promising strategy for extending the reach of bump-hunt searches at the LHC. In a weakly supervised setup, a high-quality background template is essential to produce an anomaly score but can serve a second purpose: It allows to directly estimate the background expectation in a simple cut and count setup, removing the problem of background sculpting. For imperfect background templates, the expected number of selected events from the signal region and, in particular, the corresponding systematic uncertainty needs to be well understood. I will present an improved, statistically robust determination of these quantities and discuss the resulting sensitivity of the method.

        Speaker: Mark Fanselow (RWTH Aachen University)
      • 11:10
        Kitchen Sink Anomaly Detection 20m

        Recent years have seen rapid progress in resonant anomaly detection for collider searches, but existing studies often rely on a limited set of signal benchmarks and face a trade-off between sensitive but model-dependent high-level observables and fully agnostic but less performant low-level representations. We address both limitations by introducing new simulated signal benchmarks, publicly released in a format compatible with the LHCO R&D benchmark, and by studying a broad high-level, yet highly agnostic, observable set combining Energy Flow Polynomials with subjettiness variables.
        We evaluate this combined “kitchen sink” representation against several baseline observable sets in both an idealized anomaly-detection setting and the CWoLa hunting task. Across a broad range of signal types, the combined observable set achieves the best overall sensitivity.

        Speaker: Lukas Lang (RWTH Aachen University)
      • 11:30
        Anomaly detection with normalizing flows for model-agnostic new physics searches 20m

        Discovering new particles from beyond the Standard Model remains one of the main goals of present-day particle physics. Traditional searches for new physics at the Large Hadron Collider rely on specific theoretical scenarios and simulation-based background estimates, limiting their reach and introducing modeling uncertainties. We present an anomaly detection method that uses normalizing flows to identify anomalous events without assuming a particular signal model, with the goal of estimating backgrounds directly from data rather than simulation. By training a flow on data and splitting the resulting latent representation into two independent parts, we construct two decorrelated anomaly scores that allow the background in the signal region to be estimated using the well-established ABCD method. We test this approach on a benchmark new physics scenario and show that it successfully decorrelates the anomaly scores and correctly estimates the S/√B in the signal region.

        Speaker: Rafal Maselek
      • 11:50
        Unsupervised Machine Learning for Jet-Based Anomaly Detection in ATLAS 20m

        Searches for physics beyond the Standard Model at the LHC are traditionally optimized for specific signal hypotheses. Anomaly detection provides a complementary approach by identifying events that deviate from the expected Standard Model background without relying on a particular new-physics model. In fully hadronic final states, these techniques exploit the rich information encoded in the internal structure of hadronic jets, learning directly from collision data or background-dominated samples.

        This contribution presents the current status of anomaly detection studies in fully hadronic final states within the ATLAS Collaboration. The first fully unsupervised search [1], based on a Variational Recurrent Neural Network trained directly on recorded data, is presented together with more recent developments based on Transformer architectures and Graph Neural Networks (EGAT and GIN). Particular attention is given to the different representations of jet constituents used by these models and their impact on the identification of anomalous events.

        Results obtained with the LHC Olympics benchmark dataset [2] are presented together with the first applications of these techniques to searches for high-mass diboson resonances in proton-proton collisions at √s = 13 TeV recorded by the ATLAS detector. The talk will summarize the current status of these studies and discuss their prospects for future model-independent searches at the LHC.

        References
        [1] Phys. Rev. D 108 (2023) 052009.
        [2] The LHC Olympics 2020: A Community Challenge for Anomaly Detection in High Energy Physics, Rep. Prog. Phys. 84 (2021) 124201.

        Speakers: Antonio D'Avanzo (University Federico II and INFN, Naples (IT)), Elvira Rossi (University Federico II and INFN, Naples (IT)), Francesco Cirotto (University Federico II and INFN, Naples (IT)), Francesco Conventi (Università degli studi di Napoli "Parthenope" and INFN Sezione di Napoli (IT)), Graziella Russo (University of California,Santa Cruz (US))
      • 12:10
        Look Everywhere Effects in Anomaly Detection 20m

        Machine learning–based anomaly detection can search for new physics in high-dimensional data with minimal theory bias. However, because these methods scan many possibilities at once, they suffer from a look-elsewhere effect that weakens statistical significance. We study this in weakly supervised settings and find a key trade-off: training and testing on the same data gives high sensitivity but badly miscalibrated p-values due to overfitting, while splitting the data and using a fully independent test set yields correct calibration but reduced sensitivity. Methods like early stopping help, but also cost sensitivity. We find that k-folding provides an effective balance, maintaining good calibration while preserving much of the sensitivity. Our findings are supported by numerical studies with Gaussian random variables as well as from collider physics using the LHC Olympics benchmark anomaly detection dataset.

        Speaker: Marie Hein (RWTH Aachen University)
    • 10:30 12:30
      Parallel Talks: Fast AI / Trigger
      • 10:30
        ORBIT: Online Real-time Bandwidth reduction via Information Tokens 20m

        The CMS 40 MHz Scouting program at the High-Luminosity LHC requires aggressive real-time compression of particle-flow event information to operate within stringent bandwidth constraints. We investigate the use of quantized autoencoders for converting Level-1 event representations into compact sequences of discrete tokens. We explore vector quantization, finite scalar quantization, and split-quantizer architectures that aim to separately encode morphological and amplitude information. Using the Collide-2V dataset, we evaluate these tokenization schemes in terms of compression rate, codebook utilization, constituent-level reconstruction fidelity, and preservation of jet observables including transverse momentum, mass, and substructure. We further evaluate the methods for downstream physics analyses by studying the reconstruction of the $H\rightarrow b\bar{b}$ dijet mass spectrum. Overall, these studies characterize the trade-off between token bitrate and physics fidelity, demonstrating the potential of learned tokenization as a physics-aware compression strategy for low-level particle data within the Scouting system.

        Speakers: Philipp Wagner (ETH Zürich), Younes Malek Elberkennou (ETH Zürich)
      • 10:50
        Machine learning for jets for the CMS Phase-2 Update of the L1 Trigger 20m

        At the Phase-2 Upgrade of the CMS Level-1 Trigger (L1T), particles will be reconstructed by linking charged particle tracks with clusters in the calorimeters and muon tracks from the muon stations. The 200 pileup interactions will be mitigated using primary vertex reconstruction for charged particles and a weighting for neutral particles based on the distribution of energy in a small area. Jets will be reconstructed from these pileup-subtracted particles using a fast cone algorithm. For the first time at the CMS L1T, the particle constituents of jets will be available, opening numerous opportunities to effectively leverage machine learning for multiple different purposes. This talk presents two machine learning models for tagging jets of different radii. Both follow a similar strategy of processing the jet constituents separately in a Deep Sets architecture before aggregating. For narrow jets, the jet flavor as well as the momentum are essential for designing effective trigger seeds. A Deep Sets-based architecture proves promising for both jet-flavor identification and the regression of a momentum correction factor for the jet. The model distinguishes between light-flavor jets ($uds$), gluon, $b$, and $c$ jets, as well as $\tau^+$, $\tau^-$, electron, and muon jets, going far beyond previous jet-tagging developments. For momentum regression, the model predicts individual corrections for each constituent, which are then combined to derive a correction factor for the entire jet. This approach improves upon the existing jet energy corrections, particularly in the reconstruction of invariant masses of particles. For the first time, boosted resonances will be reconstructed with dedicated large radius jet reconstruction. A neural network has been trained to identify substructure within large radius jets, enhancing triggering on boosted resonances.

        Speaker: Stella Felice Schaefer (Hamburg University (DE))
      • 11:10
        Do BitNet Gains Survive FPGA Synthesis? An Implementation-Aware Benchmark of Low-Precision Jet Classification 20m

        Real-time jet classification in high-energy physics requires high predictive performance under strict latency, throughput, and FPGA-resource constraints. Although aggressive weight quantisation promises simpler arithmetic, reductions in numerical precision and theoretical operation count do not necessarily translate into more efficient synthesised implementations. This study investigates whether the predictive gains of BitNet-style binary- and ternary-weight neural networks survive FPGA synthesis.

        The benchmark uses the public OpenML hls4ml_lhc_jets_hlf dataset, consisting of approximately 830,000 jets described by 16 high-level observables and balanced across gluon, light-quark, W, Z, and top classes. The primary task separates light-flavour QCD-like quark and gluon jets from jets labelled as W, Z, or top. Using fixed stratified data partitions and three training seeds, the study compares dense, fixed-point QKeras, HGQ, binary and ternary QKeras, BitNet binary, and BitNet-1.58 networks using two multilayer-perceptron architectures matched across model families. An unrolled XGBoost boosted decision-tree ensemble provides a tree-based reference. Neural networks are synthesised with hls4ml and decision trees with Conifer, targeting a Xilinx VU13P FPGA with a 5 ns clock period and initiation interval II=1.

        BitNet-style learned scaling consistently improves predictive performance over the corresponding binary and ternary QKeras models, with BitNet-1.58 providing the strongest performance among the tested ultra-low-bit neural networks. However, these predictive gains do not automatically produce the most efficient FPGA implementations, as the learned scaling introduces implementation-dependent overhead. Seven-bit QKeras, meanwhile, matches the dense floating-point baseline in predictive performance while reducing latency, LUT usage, and DSP usage. Among the tested implementations, HGQ provides the strongest neural resource-efficiency trade-off, and the unrolled boosted decision-tree ensemble achieves the lowest synthesised latency.

        The results demonstrate that nominal numerical precision alone is insufficient to predict FPGA efficiency. Scaling-factor realisation, structural sparsity, compiler scheduling, and model-conversion pathways can determine whether the apparent benefits of ultra-low-bit inference survive synthesis. The reported results are based on HLS C-synthesis estimates and provide a foundation for subsequent place-and-route validation.

        Speaker: Dorian Sloot (Austrian Academy of Sciences (AT))
      • 11:30
        High-performance and portable ML Inference using SOFIE 20m

        SOFIE (System for Optimized Fast Inference code Emit), being developed by the ML4EP Project at CERN, translates trained machine learning models into self-contained, low-latency C++ code that is portable, hardware-agnostic, and highly optimized while depending only on BLAS libraries.

        SOFIE achieves portability across heterogeneous computing architectures by leveraging the abstract buffer definitions provided by the alpaka[1] library. The generated code incorporates several inference optimizations, including kernel fusion, efficient memory usage, and support for quantized models, enabling high-performance inference on modern CPU and accelerator platforms.

        Beyond standalone inference, SOFIE serves as a component in several projects. It powers Yukti, a header-only interface that provides a unified zero-copy API for machine learning inference runtimes; it is integrated with RooFit as a neural surrogate for likelihood evaluation, with automatic differentiation provided by CLAD; and it could embed a hardware-agnostic, optimized compression and decompression pipeline for BOA Constrictor[2].

        In this work, we present comprehensive benchmarking and profiling results for machine learning models used in high-energy physics, including ParticleNet, ATLAS GN2, State Space Models, ParticleFlow networks, and other jet-tagging architectures. We evaluate inference latency, throughput, memory footprint, and scalability across different hardware backends, highlighting the impact of SOFIE's code generation and optimization strategies. These results demonstrate SOFIE's ability to provide portable, efficient, and high-performance inference for modern machine learning workloads in jet reconstruction, identification and other performance sensitive environments having constraints on latency and memory.

        [1] Matthes, A., Widera, R., Zenker, E., Worpitz, B., Huebl, A., & Bussmann, M. (2017, June 30). Tuning and optimization for a variety of many-core architectures without changing a single line of implementation code using the Alpaka library. Retrieved from http://arxiv.org/abs/1706.10086
        [2] Gupta, A., Doglioni, C., & Elliott, T. J. (2025). BOA Constrictor: A Mamba-based lossless compressor for High Energy Physics data. arXiv [Physics.Comp-Ph]. Retrieved from http://arxiv.org/abs/2511.11337

        Speaker: Sanjiban Sengupta (The University of Manchester (GB))
      • 11:50
        FlashJet: exact jet clustering and jet substructure reconstruction at the GPU era 20m

        Jet clustering remains one of the few steps in modern analysis and reconstruction chains that is still bound to the CPU. Machine learning workflows that wish to recluster jets inside the training loop, scan the jet radius, or access substructure dynamically must pay for repeated transfers between host and device, and reconstruction itself is steadily moving toward heterogeneous environments in which as much of the event processing as possible is expected to run on accelerators. We present flashjet, an open source library that brings the standard sequential recombination algorithms (kt, Cambridge Aachen and anti kt) natively to the GPU through custom Triton kernels. Writing the clustering in Triton, the GPU programming language of the PyTorch ecosystem, means the algorithm compiles and autotunes for the actual hardware at hand and integrates seamlessly with tensor based ML pipelines, with no separate CUDA build or hand written device code to maintain. Each kernel program clusters one event entirely on chip, the whole batch is processed in a single fused launch, and everything stays on the device. Jets, their full clustering histories, and derived substructure such as grooming, subjets and Lund plane observables all come out as tensors at negligible extra cost.
        The physics output is identical to FastJet. The implementation is validated against the FastJet reference at full numerical precision, reproducing jet momenta, groomed masses and substructure observables jet by jet on realistic simulated events, from isolated large radius jets up to the clustering of complete events with thousands of constituents. Standard analytic benchmarks of the jet substructure literature, such as the characteristic jet shapes of the anti kt algorithm and the predicted behaviour of groomed observables, are reproduced as well.
        Timing studies demonstrate the advantage of the method. Including all data movement from file to device, the Triton kernels cluster events several times faster than CPU FastJet on a single core, at the level of tens of microseconds per event on current GPUs. Exact clustering and substructure thus become cheap enough to live inside a training loop, and the library offers a concrete building block for GPU resident reconstruction, where jets can be formed and analyzed on the accelerator without ever returning to the host.

        Speaker: Chirayu Gupta (Vrije Universiteit Brussel (BE))
      • 12:10
        Application of Machine Learning to Collision Parameter Tuning in the SuperKEKB Accelerator 20m

        SuperKEKB is an electron–positron collider operating at a center-of-mass energy of 10.58 GeV and has achieved the world’s highest instantaneous luminosity. At present, collision parameters are optimized manually through a procedure known as an interaction-point (IP) knob tuning. This study aims to improve the efficiency and reproducibility of IP knob tuning, and ultimately automate the process using machine-learning techniques.
        During the 2025-2026 operation period, we actually conducted several optimization approaches: (1) simultaneous multidimensional tuning based on Bayesian optimization, (2) optimization assisted by a luminosity regression model for stability of the optimization, and (3) optimization assisted by an encoder model trained on historical machine-operation data. In this presentation, we report a comparison of these methods and discuss their applicability to automated IP knob tuning.

        Speaker: Riku Takizawa (UTokyo)
    • 12:30 14:00
      Lunch Break 1h 30m
    • 14:00 16:00
      Parallel Talks: Anomaly Detection
      • 14:00
        Domain Adaptation Against Background Sculpting 20m

        Weakly supervised anomaly detection has been shown to be an effective tool for model agnostic searches for new physics, especially in the context of resonance searches. However, to integrate weak supervision into standard resonance search analysis workflows, which rely on fits in the sidebands for the background estimation, anomaly scores need to be well behaved in the sidebands. In order to achieve this for any weakly supervised anomaly detection method, we propose a domain adaptation-based decorrelation of the anomaly detection score from the resonant mass. We demonstrate the effectiveness of this method to not only reduce background sculpting but also recover anomaly detection performance in the presence of correlated features using the LHC Olympics R\&D data set for both CATHODE and CWoLa Hunting.

        Speaker: Vincent Benne (RWTH Aachen University)
      • 14:20
        HAXAD: Anomaly Detection in the Higgs Peak 20m

        The Higgs boson, with its universal coupling to mass, provides a broadly applicable portal to sectors beyond the Standard Model and is therefore a natural anchor for anomaly detection (AD) at collider experiments. The Higgs And X Anomaly Detection (HAXAD) strategy offers a principled approach to searching for anomalies occurring in association with a Higgs boson. In the kinematic region of the Higgs peak, HAXAD employs machine-learning-based feature embedding, data- and simulation-driven background estimation, and weakly supervised classification. This contribution features the latest developments, including two new embedding strategies and a completely updated inference framework that extracts both model-independent and model-dependent limits.

        We evaluate HAXAD on an extensive simulated dataset of diphoton final states, with a large benchmark suite of signal models. HAXAD achieves strong signal sensitivity and, when benchmarked against a representative multi-category cut-based search, matches or exceeds the best individual cut-based limits for a wide variety of signal models. Together, these results establish HAXAD as a viable and compelling AD-based search strategy with novel discovery potential at colliders.

        Speaker: Dennis Noll (Stanford University)
      • 14:40
        Recent Results from Anomaly Detection Searches on CMS 20m

        In the absence of direct evidence for new physics in targeted searches, model-independent strategies are becoming increasingly important. In this talk, we present recent results of model-agnostic searches that are facilitated by advanced machine learning techniques, opening a new avenue for unbiased detection of potential new physics signals.

        Speaker: Tore Von Schwartz (Hamburg University (DE))
      • 15:00
        Towards a Statistical Interpretation of the Normalized Autoencoder 20m

        Unsupervised anomaly detection with autoencoders is a promising data-driven and model-agnostic approach for new physics searches at the LHC. However, current anomaly scores assigned by neural networks suffer from a lack of statistical interpretability. The normalized autoencoder (NAE) combines a standard bottleneck architecture with a well-defined probabilistic description. We show that the NAE ties its anomaly score to a learned likelihood via an energy-based training objective, and introduce a Bayesian version of the NAE (BNAE) that additionally provides machine-learned uncertainty estimates. We validate both on a toy model and demonstrate competitive and symmetric anomaly-tagging performance on top-versus-QCD jet tagging.

        Speaker: Jonathan Ostertag-Henning (Institute for Theoretical Physics, Heidelberg University)
      • 15:20
        Detecting and Adapting to Distribution Shift in Real-Time Anomaly Triggers 20m

        Learned anomaly triggers such as CMS AXOL1TL and CICADA are calibrated against a fixed background model, but there is no standard way to check in real time whether that calibration still holds as detector conditions change. I present two components that address this from different angles.

        The first is a sequential change-point detector that monitors trigger calibration health online. Instead of watching the output threshold alone, it tracks the anomaly score jointly with its main driving covariates (pileup, object multiplicity) through a scalar residual calibrated against expected luminosity trends. I implemented CUSUM and Bayesian Online Change-Point Detection (BOCPD) for this and compared them against standard drift detectors (ADWIN, Page-Hinkley, KSWIN).

        The second is an adaptive conformal inference (ACI) layer that recalibrates decision thresholds on the fly using miscoverage feedback from Zero-Bias control data and delayed offline validation. This matters because standard conformal prediction assumes exchangeable data, which does not hold at the LHC due to beam current decay and pileup drift. Since triggers process millions of events per second, single-event error control is not sufficient on its own, so I added online false discovery rate control (LORD, SAFFRON) evaluated over windowed batches.

        I evaluated both components on public CMS and LHC Open Data, using simulated gradual drift (pileup evolving across a fill) and abrupt shifts (subdetector conditions changing, e.g. masked readout channels). For the change-point detector, I report Average Run Length, detection latency, and false-alarm rate; for the conformal layer, I report empirical coverage, online FDR, and detection efficiency against a fixed-threshold baseline. I also examine robustness when the assumed drift model is misspecified.

        Together, these provide a model-agnostic way to detect when a trigger's background assumptions break down, and to correct for it without retraining.

        Speaker: Amishi Agrawal (KJ Somaiya School of Engineering)
      • 15:40
        Multitask, Every Region, All at Once : Fine-Tunable Region-Conditioned Simulation-Based Inference for Collider Analysis 20m

        Collider simulations simultaneously provide three sources of information: whether an event is signal or background, the physics parameters that generated each signal event, and the event kinematics from which signal regions are defined. Conventional pipelines use these sources in separate stages—training classifiers for signal discrimination, performing parameter inference for fixed analysis regions, and optimizing signal regions for specified signal hypotheses. Because all three are driven by the same simulation and the same underlying statistics, treating them as disjoint tasks leaves shared structure unexploited. We present a proof of concept in which a single region-aware model is trained jointly on all three sources. The same trained model then supports signal detection, parameter inference, and region screening as different projections of a shared conditional distribution, while providing a pretrained backbone that can be fine-tuned to a chosen region with less data than training from scratch. We validate the idea on a Gaussian bump-hunt, where all three tasks and fine-tuning behave largely as intended, and stress-test it on a non-resonant mono-jet dark-matter search that exposes its limitations. These results motivate unified learning from collider simulations as a precursor to foundation models for collider analysis.

        Speaker: Yong Sheng Koay
    • 14:00 16:00
      Parallel Talks: Simulation Based Inference
      • 14:00
        Proton Structure from Neural Simulation-Based Inference at the LHC 20m

        The precise determination of the parton distribution functions (PDFs) of the proton is an essential ingredient for LHC analyses, including for those at the upcoming High-Luminosity LHC. So far, PDFs are determined from global fits to binned low-dimensional data obtained from unfolded hard-scattering cross section measurements. In this talk, we demonstrate the feasibility of neural simulation-based inference (NSBI) to constrain the proton PDFs using a high-dimensional unbinned data set. As a proof-of-concept, we determine the gluon PDF from simulated data of top quark pair production at the LHC with $\sqrt{s} = 13$ TeV. Taking into account both experimental and theoretical systematic uncertainties in the detector-level features, we demonstrate how the NSBI pipeline achieves significant improvements in precision compared to existing low-dimensional binned analyses.

        Speaker: Jaco ter Hoeve (The University of Edinburgh)
      • 14:20
        Reliable Likelihood Ratios for Neural Simulation Based Inference in HEP 20m

        Neural Simulation Based Inference (NSBI) is a collection of statistical machine learning methods that learn the likelihood or posterior based on high-dimensional input data. One flavor of NSBI, Neural Likelihood Ratio Estimation (NLRE), has emerged as a powerful method in the domain of High Energy Physics (HEP), having recently been used in Higgs-boson measurement [1] at the Large Hadron Collider (LHC). NLRE uses the likelihood ratio trick [2] to transform the Bayes optimal binary classifier to a ratio of probability densities, as a surrogate for the ratio that can be used directly in a per-event maximum likelihood process to perform unbinned measurements. It has the inherent advantage over binned template histogram fits as it doesn't incur loss of information by forced dimensionality reduction nor the aggregation process of binning. However, due to the direct usage of neural surrogates, NLRE requires highly reliable classifiers. This fact can prevent the utilization of more powerful, deep networks, which are known to suffer from calibration shortcomings [3]. Leveraging the core infrastructural element of the NEEDLE project [4], needle-sbi, an efficient workflow orchestrator for large-scale deployment and tracking of neural surrogates enabling practicable studying of NSBI, we investigate this problem from two perspectives: improving the estimator and restructuring the question posed to the model.

        We present an extensive validation and calibration suite for NLRE classifiers, integrated into, and automated by needle-sbi, with which we demonstrate the insufficiency of common binary classifier post-hoc calibration approaches for improving estimated probability ratios in the NLRE setting. Those methods, like temperature scaling, isotonic regression, and Platt scaling rely [5] on remapping classifier scores, but no function of the score alone can repair a residual that differs between events sharing the same score. Therefore, we further investigate more expressive methods such as modular model extension [6] acting on the learned representation and adapter-based fine-tuning [7]. Studies use transformer models trained on FAIR Universe Higgs uncertainty challenge dataset [8] as a proxy for realistic HEP data analyses, and a custom three-body decay model [9] for its tractable likelihood and generator-level densities, the latter of which we also investigate as a direct augmentation of the training objective.

        Restructuring the question begins from what NLRE does well: classifiers are simple and not too computationally costly to train. However, HEP statistical workflows frequently require combining multiple neural probability ratio surrogates which further compounds small classifier errors. Neural Likelihood Estimation (NLE) directly produces absolute densities that don't require composition, but solver-based NLE with diffusion models [10] or flow matching [11] is much more computationally expensive. We introduce Density Ratio Diffusion Probabilistic Models (DRDPM) that hybridize the two approaches: they deliver NLE, while retaining classifier-training simplicity of NLRE. DRDPM hinges on a multi-class classifier trained to discern the discrete timesteps of a forward diffusion process as it interpolates between the data distribution and a standard Gaussian anchor, from which conditional log-likelihood is obtained as the difference of two boundary logits, combined with the closed-form Gaussian term. Demonstrated on the three-body decay model, DRDPM yields absolute likelihoods at the cost of a single forward pass, increasing the speed of NLE inferencing by an order of magnitude or more compared to solver-based NLE.

        Because DRDPM relies on differences between boundary-class outputs, its calibration becomes the central methodological challenge and a natural direction for future work merging these two research tracks.

        References:

        [1] The ATLAS Collaboration, Measurement of off-shell Higgs boson production in the $H^*\rightarrow ZZ\rightarrow 4\ell$ decay channel using a neural simulation-based inference technique in 13TeV pp collisions with the ATLAS detector, Reports on Progress in Physics, vol. 88, p. 057803, 2025.
        [2] J. Brehmer, K. Cranmer, G. Louppe, and J. Pavez, A guide to constraining effective field theories
        with machine learning
        , Phys. Rev. D, vol. 98, p. 052004, 2018.
        [3] C. Guo, G. Pleiss, Y. Sun, and K. Weinberger, On calibration of modern neural networks, ICML, 2017.
        [4] NEEDLE open-source release version GitHub: https://github.com/needle-sbi/needle-sbi/
        [5] T. Silva Filho, H. Song, M. Perello-Nieto, R. Santos-Rodriguez, M. Kull, and P. Flach, Classifier calibration: a survey on how to assess and improve predicted class probabilities, Machine Learning, vol. 112, pp. 3211--3260, 2023.
        [6] A. Rahimi, T. Mensink, K. Gupta, T. Ajanthan, C. Sminchisescu, and R. Hartley, Post-hoc calibration of neural networks by g-layers, arXiv:2006.12807, 2022.
        [7] E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen, LoRA: Low-rank adaptation of large language models, ICLR, 2022.
        [8] FAIR universe - HiggsML Uncertainty Challenge: https://www.codabench.org/competitions/2977/
        [9] Three-Body Decay Monte Carlo package: https://sjiggins.github.io/VegasThreeBodyDecay-docs/
        [10] Y. Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole, Score-based generative modeling through stochastic differential equations, ICLR, 2021.
        [11] Y. Lipman, R. T. Q. Chen, H. Ben-Hamu, M. Nickel, and M. Le, Flow matching for generative modeling, ICLR, 2023.

        Speaker: Nino Kovacic (University of Zagreb (HR))
      • 14:40
        Iterative Simulation-Based Inference for the LHC 20m

        Simulation-based inference (SBI) is a powerful tool for likelihood-free parameter estimation, but often requires large numbers of simulated events. At the LHC, where event generation can be computationally expensive, particularly with restrictive generator-level cuts and detector simulation, this can become a major bottleneck. We demonstrate sequential SBI for the inference of resonance masses and widths, showing that it can substantially improve sample efficiency by adaptively concentrating simulations in regions of parameter space favored by the data. This enables accurate inference with significantly fewer simulated events, making SBI more practical for computationally demanding LHC analyses.

        Speaker: Ranit Das (Heidelberg University)
      • 15:00
        Probing the Higgs Charge Parity at the LHC with Invertible Neural Networks 20m

        CP-sensitive observables in H+2jet production, such as the azimuthal angle between the leading jets, are defined at parton level but measured at detector level, requiring unfolding to recover them. We present a conditional invertible neural network (cINN) that learns a full posterior over parton-level kinematics given detector-level observables trained jointly across multiple EFT coupling scenarios. We discuss design choices driven by the CP-sensitive target observable and progress toward a general, data-ready unfolding pipeline.

        Speaker: Nityaansh Parekh (Michigan State University (US))
      • 15:20
        Trusting Generative Unfolding 20m

        Machine learning enables unbinned unfolding with per-event posteriors, but a measurement is only as good as its error bars. We dissect the uncertainty budget of conditional flow matching unfolding for WZ production. Bootstraps propagate the statistics of the training sample and, by reweighting, of the data. We show how weight sampling from a Bayesian neural network relates to the spread of retraining ensembles. This turns generative unfolding into a measurement tool with a transparent error budget.

        Speaker: Sebastian Pitz (LPNHE Paris)
      • 15:40
        Unfolding without Iterations, Adversaries, or Surrogates 20m

        Correcting measurements for detector effects is a pressing inverse problem in LHC physics. Current methods solve this problem by relying on iterative refinement, minimax optimization, or a surrogate forward mapping. In this talk, I present Adversary-free Unfolding SanS Iteration or Emulation (AUSSIE), which dispenses with these mechanisms while remaining asymptotically correct. AUSSIE unfolds by reweighting a reference simulator in similar fashion to OmniFold. However, its new kernel-based loss function yields one-shot solutions with minimal bias toward the reference distribution. I show results for AUSSIE applied to a range of unfolding tasks, from low-dimensional examples to full-phase-space jet substructure.

        Speaker: Ayodele Ore
    • 16:00 16:30
      Coffee Break 30m
    • 16:30 17:30
      Parallel Talks: Scaling Laws
      • 16:30
        Scaling laws for amplitude surrogates 20m

        Fast and precise evaluations of scattering amplitudes even in the case of precision calculations is essential for event generation tools at the HL-LHC. We explore the scaling behavior of the achievable precision of neural networks in this regression problem for multiple architectures, including a Lorentz symmetry aware multilayer perceptron and a fully Lorentz equivariant transformer using Lorentz Local Canonicalization (LLoCa). This study addresses in particular the scaling behavior of uncertainty estimations using state of the art methods.

        Speaker: Joaquin Iturriza Ramirez (LPNHE - Sorbonne Université)
      • 16:50
        Neural Scaling Laws for Jet Generation 20m

        Scaling laws have become a central topic in modern machine learning, providing a quantitative understanding of how model performance improves with increasing model size, training data, and compute. They also offer insights into whether learning is approaching the information limits of a given dataset. In this contribution, we present the first study of neural scaling laws for generative jet modeling. Beyond the conventional next-token prediction validation loss, we also investigate the scaling behavior of the sliced Wasserstein distance computed for five high-level jet observables, providing a physics-motivated measure of generative performance. As a function of model size, we observe that both metrics exhibit the expected logarithmic scaling behavior. In contrast, scaling with dataset size and training compute is significantly weaker for both the validation loss and the sliced Wasserstein distance. We analyze this behavior by introducing the concept of a learnable window, and argue that autoregressive next token prediction on jet constituents exhibits comparatively rapid saturation relative to language-model studies. We discuss several possible explanations for this observation, including the intrinsic stochasticity of QCD jet formation and the fundamental differences between supervised prediction and generative modeling.

        Speaker: Anna Kindsvater (University of Hamburg)
      • 17:10
        Scaling Laws for the Unified Particle Transformer in Heavy-Flavour Tagging 20m

        Transformer-based taggers have become the workhorse of heavy-flavour identification at the LHC, with the Unified Particle Transformer (UParT) folding flavour classification, track-level auxiliary tasks and regression into a single architecture. Their development has nonetheless remained largely empirical: model capacity, training-sample size and input granularity are chosen by convention rather than derived from a predictive framework. We present a systematic study of the scaling behaviour of UParT-style taggers. Varying model size and training statistics over a wide range, we fit the tagging loss to the parameterization L(N, D) = E + A/N^α + B/D^β and observe clean power-law scaling in both quantities, together with a clearly non-zero irreducible term. We argue that this floor is physical rather than architectural and interpret it in terms of the achievable light-jet rejection at fixed b-tagging efficiency and other flavour tagging metrics.

        Beyond the individual power laws, we study where each configuration saturates and extract the compute-optimal allocation between parameters and training jets, in direct analogy to the Chinchilla token-to-parameter analysis for language models. Strikingly, the optimal jet-to-parameter ratio follows a very similar trend, suggesting that the compute-optimal balance is governed by the structure of the training objective rather than by the specifics of the data domain. We further quantify how loss improvements propagate into misidentification rates and how the resulting Pareto front interacts with the inference-latency budgets of offline reconstruction and the high-level trigger.

        Speaker: Pavlo Kashko (Vrije Universiteit Brussel (BE))
    • 16:30 17:30
      Parallel Talks: Simulation Based Inference
      • 16:30
        A fully-differential charged particle phase space measurement of $Z$+jets production in $pp$ collisions at $\sqrt{s} = 13$ TeV with the ATLAS detector 20m

        The production of a $Z$ boson decaying to muons in association with hadronic jets is a benchmark process in proton--proton collisions at the Large Hadron Collider. It provides a clean experimental signature with high selection efficiency and purity while probing a wide range of Standard Model dynamics. Measurements of this process are typically reported as binned differential cross sections in a limited set of pre-defined observables, which reduces the information retained and hinders reinterpretation. In this talk, the OmniFold unfolding method, which utilizes iteratively trained transformer neural networks, is used to produce an unbinned and highly-differential measurement of the $Z$+jets cross section using 140~fb$^{-1}$ of $\sqrt{s} = 13$ TeV proton--proton collision data collected by the ATLAS detector. This constitutes the first published unbinned, fully differential measurement of a collider cross section in the charged-particle phase space within a fiducial volume. The talk covers details of the application of Omnifold to data, including the novel use of pretrained representations to improve unfolding accuracy and reduce computational cost. Several applications of the measurement to constrain novel jet substructure observables will also be discussed.

        Speaker: Kevin Thomas Greif (University of California Irvine (US))
      • 16:50
        Learning the latent structure of background distributions 20m

        Topic modeling techniques allow to learn robust representations of a data in terms of latent themes or topics. In this talk, I will present an application of Latent Dirichlet Allocation that exploits information from multiple Monte Carlo simulation setups of known processes to learn the shapes of observable distributions directly from data. This approach provides a general framework to infer the latent probabilistic structure underlying the observed data, which permits flexible and data-driven background inference and uncertainty estimation. I will demonstrate this method in a specific application to multijet final states at collider experiments.

        Speaker: Santiago Tanco
      • 17:10
        Exploring data-driven likelihood estimation with normalizing flows for CRESST 20m

        CRESST (Cryogenic Rare Event Search with Superconducting Thermometers) is a direct dark matter detection experiment located at the Laboratori Nazionali del Gran Sasso (LNGS) in Italy. It searches for dark matter–nucleus interactions using scintillating cryogenic calorimeters, pushing its energy threshold ever lower to improve sensitivity to low‑mass dark matter. At these thresholds, however, the analysis is complicated by an exponentially rising event rate below 200 eV, referred to as the low‑energy excess, whose origin remains unknown. This unexplained background challenges traditional, parametrized likelihood models of the data. In this talk, we explore the use of normalizing flows to learn the underlying probability density of the CRESST data directly, and discuss how such data‑driven density estimators can provide an alternative likelihood construction and potentially a more flexible framework for setting sensitivity limits.

        Speaker: Danae Danielle Valdenaire (Austrian Academy of Sciences (AT))
    • 19:00 22:00
      Dinner 3h
    • 09:00 10:00
      Plenary Presentations
      • 09:00
        GPU-deployment in online reconstruction at ALICE 30m

        TBA

        Speaker: Felix Weiglhofer (CERN)
      • 09:30
        Machine Learning for Hadronization 30m

        Hadronization, the transition between unobservable quarks and gluons to observable hadrons, is a key aspect of the theoretical framework of particle physics. However, it is a fundamentally challenging process due to its non-pertubative nature, and thus event generators implement empirical models based on QCD insights. In this talk, I'll detail how Machine Learning has been incorporated into different aspects of hadronization modeling, including parameter tuning uncertainty quantification, and ML-based models, with an emphasis on ongoing challenges and possible ways forward.

        Speaker: Manuel Szewc
    • 10:00 10:30
      Coffee Break 30m
    • 10:30 12:30
      Parallel Talks: Equivariant NN
      • 10:30
        Virtues and Vices of Equivariant Transformers 20m

        We study for the first time the benefit of Lorentz-equivariant transformers for large-size jet tagging and flavor tagging. To control their computing demands, we optimize their implementations for inference cost metrics. In our scaling studies, we find that Lorentz-equivariant networks outperform standard transformers provided geometric features are relevant. This holds true in an idealized world as well as for limited resources. The limited gain from Lorentz-equivariance provides interesting input to the development of foundation models for LHC data.

        Speaker: Jonas Spinner (Durham University)
      • 10:50
        Economical Jet Taggers -- Equivariant, Slim and Quantized 20m

        Modern machine learning is transforming jet tagging at the LHC, but the leading transformer architectures are large, not particularly fast, and training-intensive. We present a slim version of the L-GATr tagger, reduce the number of parameters of jet-tagging transformers, and quantize them. We compare different quantization methods for standard and Lorentz-equivariant transformers and estimate their gains in resource efficiency. We find an order-of-magnitude reduction in energy cost for an moderate performance decrease, down to 1000-parameter taggers. This might be a step towards trigger-level jet tagging with small and quantized versions of the leading equivariant transformer architectures.

        Speaker: Antoine Petitjean (Heidelberg University)
      • 11:10
        Group equivariance in infrared and collinear safe graph neural networks for jet classification 20m

        Improving interpretability is essential for building robust
        and trustworthy machine learning tools in collider physics. To address this challenge, we systematically investigate equivariant and IRC-safe graph neural networks for jet classification. Using simulated jet datasets, we compare IRC-safe architectures with inbuilt E(2) and O(2) equivariance in the rapidity-azimuth plane against IRC-safe and -unsafe baselines in terms of classification performance, robustness to soft emissions, and latent representation structures. Our analysis shows that IRC-safe and symmetry-aware networks are more stable across training instances and distribute their latent variance across multiple interpretable directions. By regressing Energy Flow Polynomials onto the leading principal components, we establish a direct correspondence between learned representations and established IRC-safe jet observables. These results demonstrate that embedding symmetry and safety constraints not only improves robustness but also grounds network representations in known QCD structures.

        Speaker: Vishal Singh Ngairangbam
      • 11:30
        Enforcing IRC Safety and Lipschitz-Constrained Insensitivity to Non-Perturbative Effects in Transformers 20m

        IRC safety has long guided the design of robust jet substructure observables, and this principle has increasingly been built into machine-learned taggers as well. For attention-based architectures such as the Particle Transformer (ParT), however, it is non-trivial to enforce IRC safety without discarding the pairwise, energy-dependent features that make these models powerful. In this work, we show that IRC safety can be built directly into the attention mechanism by removing energy information from per-particle tokens and pairwise features, and instead injecting it as an additive $\log E_i + \log E_j$ bias on the pre-softmax attention logits. We prove that the resulting attention output for "ParT-IRC" is invariant under soft emissions and collinear splittings.
        Building on this, we further construct a Lipschitz-constrained variant, "L-ParT-IRC", which combines our IRC-safe attention with modified self-attention and spectral normalization to bound the network's sensitivity to non-perturbative corrections. We benchmark ParT-IRC against the standard ParT on quark/gluon tagging, finding comparable in-distribution performance together with improved robustness under a Pythia-to-Herwig out-of-distribution generalization test, and we discuss the tradeoffs introduced by the Lipschitz constraint. We also demonstrate that L-ParT-IRC outperforms the Lipschitz Energy Flow Network (L-EFN) while remaining similarly insensitive to non-perturbative effects.

        Speaker: Umar Sohail Qureshi (Vanderbilt University)
      • 11:50
        One Generator, Any Process: LLM-Conditioning for the LHC 20m

        Neural network training for LHC event generation should, ideally, benefit from common high-level patterns in different processes. We propose novel conditioning schemes for continuous parameters, process labels, and Feynman diagrams. We employ pre-trained LLMs as multi-modal foundation models to provide descriptive embeddings for an autoregressive transformer. With such high-level physics-inductive bias the generative networks converge faster, provide better result, and generalize to unseen processes.

        Speaker: Thanush Sivagnanalingam (University Heidelberg)
    • 10:30 12:30
      Parallel Talks: Quantum Machine Learning
      • 10:30
        Neural-network quantum states for exotic hadrons and nuclei 20m

        Neural-network quantum states provide a flexible representation of high-dimensional many-body wave functions, offering a promising approach to quantum systems that remain challenging for conventional numerical methods. In this talk, I will present their applications to exotic hadrons and nuclei. By combining expressive neural-network wave functions with variational Monte Carlo and incorporating the relevant physical symmetries, we solve the full many-body problem for multiquark systems and quarkonium–nucleus bound states. These calculations provide quantitative predictions for the exotic hadron spectra, including states closely related to ongoing and future searches at the LHC. Our results demonstrate that neural-network quantum states constitute a powerful and broadly applicable computational framework for investigating the structure and spectroscopy of exotic hadronic and nuclear systems.
        Refs: Phys.Rev.Lett. 136 (2026) 7, 071901; arXiv 2606.09254

        Speaker: weilin wu (Peking University)
      • 10:50
        (ZOOM) Variational Quantum Classification Heads for Particle Transformer-Based in Jet Tagging 20m

        In collider physics, jet tagging is a key classification task in which models must identify the initiating particle from a reconstructed jet's internal structure. Though their final predictions are often produced by classical neural-network heads, Particle Transformer architectures achieve strong performance by learning interactions among jet constituents. In this paper, the usefulness of a variational quantum circuit as an alternate decision layer for Particle Transformer representations is investigated.

        A pretrained Particle Transformer serves as a frozen feature extractor using a controlled subset of public JetClass dataset. The resulting jet representations can be encoded into simulated four-to eight-qubit variational quantum circuits by compressing them into low-dimensional vectors. Linear, parameter-matched, and deeper multilayer-perceptron classifiers trained on the identical compressed inputs are compared to the quantum heads. Additionally, a classical classifier using the original, uncompressed transformer representation is included, allowing the performance of the quantum model itself to be distinguished from the information loss resulting from compression.

        Binary and selected multiclass jet-tagging are taken into consideration in this work. Classification accuracy, ROC-AUC, background rejection, calibration, stability across random initializations, trainable parameter count, and simulator-based computing cost are used to evaluate performance. The bottleneck dimension, circuit depth, entanglement pattern, and training-set size are varied in further tests. The goal is to provide a assessment of where variational quantum classification heads may be competitive and where current limitations still exist, rather than assuming that the quantum model will outperform classical alternatives.

        Speaker: Nadia Sharna (Bursa Technical University)
      • 11:10
        Normalising Flow-Assisted Neural Quantum States for Ground-State Estimation 20m

        Neural quantum states provide expressive variational representations of quantum many-body wavefunctions.
        However, their practical performance depends on how well they can sample the relevant configurations from
        an exponentially large Hilbert space. Conventional Markov chain Monte Carlo methods can mix slowly between separated high-probability regions, particularly in frustrated and strongly correlated systems.
        We introduce a hybrid variational framework in which a continuous normalising flow learns an auxiliary
        sampling distribution over a discretised effective subspace, while an independent variational ansatz learns the
        wavefunction amplitudes. This separates the task of finding the relevant support from the task of estimating
        the amplitudes within it. The normalising flow can therefore explore non-local regions of configuration space
        without relying on a sequence of local updates.
        We apply the method to the square-lattice J1–J2 Heisenberg model and compare it with conventional Metropolis sampling using matched variational ansätze and optimisation settings. Our results show that flow-assisted
        sampling remains competitive across the system sizes studied and improves variational ground-state estimation in several regimes. This suggests that generative-model-assisted sampling can provide a useful practical
        approach to quantum many-body simulation.

        Speaker: Timur Sypchenko (IPPP)
      • 11:30
        Continuous-Variable Photonic Quantum Machine Learning for High-Energy Physics: The 1-Particle 1-Qumode Framework 20m

        As next-generation colliders reach unprecedented energies and luminosities, novel computing techniques become essential to meet the resulting computational challenges. One area of promise is quantum machine learning (QML), which combines the quantum effects of superposition and entanglement with classical optimisation techniques. Within QML, photonic devices are of particular interest due to their natural compatibility with continuous data and the highly non-trivial inter-qumode correlations they can induce. Here we introduce the 1P1Qm encoding scheme, the continuous-variable analogue of the qubit-based 1P1Q framework, in which classical jet constituents are encoded onto quantum circuits using one particle per qumode. Under this scheme, we study quantum autoencoders for anomaly detection and supervised quantum classifiers for signal/background discrimination, benchmarking both against classical reference models on a $t\bar{t}$ and $Z(\rightarrow \nu\nu)$+jets dataset. In both cases, the quantum models outperform their classical counterparts in the low-data regime and remain competitive at larger training-set sizes, while using substantially fewer trainable parameters.

        Speaker: Louis Choron (Imperial College (GB))
      • 11:50
        Information-theoretic and quantum-informed approaches for jets at the LHC 20m

        Jet substructure studies at the LHC often rely on observables that do not fully capture all possible correlations between constituents, especially at higher irreducible orders. We present two approaches for modelling this structure, based on information theory and quantum information geometry. Both methods use the Fisher Information Matrix, the first being constructed classically using the statistics of the input variables, and the second from the quantum metric tensor, also referred to as the Quantum Fisher Information (QFI) matrix.
        In the first approach, we start from classical correlator observables such as the energy correlator functions (ECFs) and energy-energy correlators (EECs), though the approach is general and can be extended to observables such as Energy Flow Polynomials (EFPs) which form a complete basis. Pairwise Fisher graphs, built from the covariances of these observables, cannot distinguish an irreducible multi-observable radiation pattern from a collection of ordinary pairwise correlations. We then show that these irreducible correlations can be better understood using higher-order extensions of the Fisher Information, namely the Amari-Chentsov tensor at third order, and higher order cumulant tensors thereafter. We establish a Fisher-correlator-hypergraph triality by showing that the same tensor (at arbitrary order) constructed from a jet observable basis can serve as a coefficient in a local Kullback-Leibler (KL) expansion, as a connected cumulant of the correlator observables, and as a signed hyperedge weight linking a hypergraph built from these observables. This relation allows for a physics-informed construction of hypergraphs from measured or simulated jet observables (EECs, ECFs or EFPs), supplies weights for higher-order graph Laplacians, and provides a criterion for the compression of observable bases while retaining irreducible higher-order information. We demonstrate the applications of this relation on a set of simple tasks: jet tagging using a BDT-based classifier, and a low-capacity message-passing graph neural network (GNN). The resulting hypergraph-based approaches are shown to retain higher-order structure better than pairwise graphs, and provide a useful inductive bias for designing machine learning algorithms to learn from jet substructure observables.

        In the second approach, we use the QFI matrix as a complementary representation of kinematic input data, to improve jet tagging performance of large machine learning models such as GNNs and transformers. This matrix is extracted from a variational quantum circuit (VQC) trained for the same task, and serves as a representation of the intrinsic geometry of the circuit’s state manifold. This can represent structure that classical feature engineering alone finds difficult to learn. With this construction, we fuse these additional input features into an existing classical architecture, and show that this quantum-informed enhancement of classical-only inputs leads to noticeable improvements in jet tagging performance for both a simple GNN architecture, and the state-of-the-art Particle Transformer (ParT) model. These gains are measured using the AUC score and the background rejection, and remain statistically significant across multiple random seeds. We therefore demonstrate that quantum-geometric information can be extracted from classically simulated quantum circuits and be used to improve the performance of large ML models even in the NISQ era, before fault-tolerant quantum hardware becomes available.

        Speaker: Aritra Bal (KIT - Karlsruhe Institute of Technology (DE))
      • 12:10
        Lund Plane to Bloch (LP2B) encoding for object and polarization tagging with quantum jet substructure 20m

        The application of quantum algorithms to jet substructure analysis is of growing interest as Noisy Intermediate-Scale Quantum (NISQ) hardware continues to mature in qubit count and gate depth. Jet substructure remains essential for addressing challenges at the LHC and beyond, notably object classification and polarization tagging. However, existing quantum machine learning approaches typically rely on data representations that suffer from infrared and collinear (IRC) unsafety, sensitivity to non-perturbative effects, or poor scalability.
        In this talk, we introduce the Lund Plane to Bloch (LP2B) [1] encoding, which maps a theoretically clean and robust representation of jet kinematics directly into qubit states. Leveraging this encoding, we implement a Quantum Tree-Topology Network (QTTN) that natively embeds the hierarchical structure of the Lund tree. We evaluate the QTTN across multiple benchmarks, comparing it with classical machine learning architectures and the standard "one particle - one qubit" (1P1Q) encoding on polarization, W boson, and top quark tagging tasks, including in the low-data regime. The results show that, despite its low parameter count, the QTTN achieves competitive performance with classical baselines, demonstrates enhanced sensitivity compared to the 1P1Q encoding, improves the performance-to-cost trade-off, and exhibits enhanced stability in low-data regime and reduced sensitivity to generator-specific parton shower and hadronization models. Finally, the QTTN is validated on real quantum hardware using a 3-qubit solid-state NMR SpinQ device.

        [1] https://link.springer.com/article/10.1140/epjc/s10052-026-16142-9

        Speaker: Luca Della Penna (Universita e INFN, Perugia (IT))
    • 12:30 14:00
      Lunch Break 1h 30m
    • 14:00 15:00
      Parallel Talks
      • 14:00
        4D Particle Tracking in the NA62 GigaTracker with Transformer-Based Architectures 20m

        Accurate particle tracking using the GigaTracker (GTK) silicon pixel detector represents
        a mission-critical stage in the data processing pipeline of the NA62 experiment at CERN,
        which is dedicated to the precision measurement of the ultra-rare decay K+ → π+ν ¯ν.
        Operating in a high-intensity environment with a beam rate of up to 750 MHz, the GTK
        provides the momentum and direction measurement as well as timing information for the
        incoming beam particles. Traditional tracking approaches, relying on local combinatorial
        algorithms, suffer from intrinsic limitations due to high pile-up conditions and the resulting
        combinatorial background, significantly increasing the rate of fake tracks. In this work, we
        present the first Transformer-based reconstruction algorithm developed specifically for the
        NA62 GTK, designed to exploit the remarkable single hit time resolution of the detector
        of O(100 ps). Formulating the tracking challenge within an edge classification framework,
        the architecture employs a Transformer encoder to generate rich, global embeddings of each
        detector hit’s features. These representations are subsequently used to compute connectivity
        scores between admissible hits across consecutive stations. The models have been trained
        and extensively validated on high-fidelity Monte Carlo simulation samples, demonstrating
        excellent generalization capabilities and robustness across varying beam intensities and data-
        taking periods. The results show a sharp reduction in the number of false-positive tracks
        and a substantial increase in purity while maintaining a tracking efficiency comparable
        to or exceeding the standard algorithm. The architecture is currently used as the default
        reconstruction method in the NA62 C++ software framework. The improved reconstruction
        directly translates into enhanced performance for the downstream K − π matching task,
        which is central to background rejection in the experiment.

        Speaker: Gemma Tinti (INFN e Laboratori Nazionali di Frascati (IT))
    • 14:00 15:00
      Parallel Talks: Large Language Models
      • 14:00
        Fine-tuning Small Language Models for Particle-Physics Lagrangian Generation 20m

        Recent advances in language models have sparked growing interest in their use for specialized scientific applications. In this work, we apply supervised fine-tuning to adapt a small pretrained language model to a task in theoretical particle physics. We consider the generation of particle-physics Lagrangians from structured field descriptions. Given a set of fields, the model must generate the corresponding Lagrangian with the required objects and contractions.

        On a separate evaluation dataset, more than 95% of the generated Lagrangians fully matched the reference expressions in terms of the required objects and contractions. Additional experiments with different model sizes and prompting strategies showed that increasing the number of model parameters did not necessarily improve performance and that the results were sensitive to how the task was presented in the prompt.

        Our results show that a language model with three billion parameters can accurately learn the mapping from field descriptions to the corresponding Lagrangians. Whether the model captures general physical principles or primarily exploits statistical regularities in the chosen representation remains an open question.

        Speaker: Paul Schmidt
      • 14:20
        Classification-Generative Supervised Fine-Tuning for Hierarchical Tagging of ML4HEP Literature 20m

        We propose CG-SFT, a parameter-efficient framework that turns a causal large language model into a structured hierarchical multi-label classifier without adding a classification head. The label space is encoded as a fixed sequence where each label is followed by a <yes>/<no> decision token. Supervision is applied only to these decision positions using cross-entropy, binary cross-entropy, and Dice losses, plus a hierarchy penalty. New decision tokens are initialized with lightweight adapters and combined with LoRA updates on Qwen2.5-7B. We evaluate on 1,819 ML4HEP paper abstracts annotated with a 91‑label taxonomy (77 parent–child relations, average 2.8 positives per paper). Exact‑match accuracy (all 91 decisions correct) is used as the primary metric. On the held‑out test set (186 papers), CG‑SFT achieves 60.2% exact match, compared to 1.1% for the unfine‑tuned backbone and 21.0% for GPT‑5.4 Max—gains of 59.1 and 39.2 percentage points, respectively. Training takes about one hour on a single A800 GPU. Our approach bridges generative fine‑tuning and discriminative classification, enabling scalable curation and retrieval of scientific literature with sparse hierarchical labels. Furthermore, the curated data can be distilled into design knowledge for ML models in particle physics, laying the groundwork for future agent-assisted model development.

        Speaker: Ziyue Chen (IHEP)
      • 14:40
        Steering in Theory Space: Representation Engineering of a Lagrangian Transformer 20m

        Foundation models embed not only their training examples but the space those examples were drawn from. A transformer trained to write Lagrangians symbolically (such as BART-L) thus embeds the space of theories itself. In this work, we investigate the navigation of this learned theory space using activation steering, adapted from interpretability work on LLMs. Using only small contrastive datasets, we build steering vectors that take us from one theory domain to another. This hence reframes theory construction as a navigation problem, where reaching a theory of interest becomes a matter of finding and following directions, rather than an exhaustive scan of a landscape. We demonstrate navigations such as from anomalous to anomaly-free field content, from a base model to its supersymmetric and gauge-extended counterparts, and from charge-violating to charge-conserving Lagrangians. All of this requires no retraining or finetuning at all. The method also serves its original purpose of interpretability. By testing which operators can be recovered as vectors, we can ask whether the learned theory space is sufficient. We find directions corresponding to several Lorentz and gauge representation changes, including one that maps scalars onto fermions in the manner of a supersymmetry generator, as well as failure modes we associate with the limited size of our model and dataset. These results indicate that the model has, to some degree, embedded a navigable space of theories, and point toward a novel approach to theory search which is directed rather than exhaustive.

        Speaker: Yong Sheng Koay
    • 15:00 15:30
      Organisatoricals: Closing Remarks
      Convener: Dr Claudius Krause (MBI Vienna (ÖAW))