Speaker
Description
Experimental high energy physics analysis is traditionally a multi-year effort dominated by repetitive implementation that demands little physics insight. LLM-based AI agents can now autonomously execute substantial portions of this pipeline, collapsing the implementation bottleneck to roughly ten hours of wall-clock time. This talk surveys the rapidly developing landscape of agentic AI in high energy physics, anchored by Just Furnish Context (JFC), a framework built on Claude Code that integrates autonomous analysis agents, literature-based knowledge retrieval, and a multi-agent review system mirroring the tiered review structure of a real collaboration. Given only a short physics prompt and a dataset, JFC plans and executes the full chain: event selection, background estimation, systematic uncertainty quantification, blinded statistical inference, and publication-grade documentation, with human oversight concentrated at a single unblinding gate. Demonstrations on ALEPH, DELPHI, and CMS open data include a CMS Run-1 $H\to\tau\tau$ measurement consistent with the published value and the first primary Lund jet plane density measured in $e+e-$ collisions, to our knowledge the first novel HEP measurement produced autonomously by an AI agent.
I will situate these results among other recent agentic efforts across the community and confront the question these systems force: how do we know an agent's analysis is correct? I will present ongoing benchmarking work evaluating the robustness and physics fidelity of agentic analyses across models and tasks, and argue that as implementation cost collapses, the bottleneck shifts from coding capacity to physics ideas, review bandwidth, and trust.