Speaker
Description
Mechanistic interpretability aspires to reverse engineer AI models by breaking down black-box weights and activations into human-understandable features and circuits. A leading approach for mechanistic interpretability is sparse dictionary learning, which trains an encoder to map activations into a sparse code and a decoder to reconstruct them. The corresponding architecture is called sparse-autoencoder (SAE), and it has been shown to uncover safety-relevant concepts such as deception, bias, and harmful content, enabling targeted interventions on model behavior. In this talk, we will go through applications of SAEs in language and vision models, as well as present new SAE variants based on tensor decompositions that enhance expressivity at a fraction of the parameter cost.