Speaker
Description
A growing body of work on large language models has focused on steering vectors, which are linear directions within a model’s latent space that correspond to interpretable language concepts the model has learned. However whether such interpretable directions exist in models trained on physics data, that correspond to physics concepts, remains underexplored. We investigate this in the context of jet tagging, on a transformer architecture with no physics-motivated inductive biases trained on the JetClass dataset. We find that such steering vectors do seem to exist and we show specific directions that correspond to physically meaningful concepts like jet mass, constituent particle multiplicity, and jet momentum. We provide both correlational evidence through linear probing and causal evidence through steering vector interventions, demonstrating that these directions influence model predictions. For example, applying the constituent particle multiplicity direction as a steering vector to a batch of Hcc jets shifts the model's prediction toward Hgg. We also touch on broader potential applications and characteristics of steering vectors within jet tagging models, such as whether steering vectors can be used to calibrate synthetic data to better mimic observed data and whether these linear directions differ as you add inductive biases to the model.