28 September 2026 to 2 October 2026
Sun Hall Hotel
Asia/Nicosia timezone

New directions in interpreting AI models

1 Oct 2026, 12:00
30m
Sun Hall Hotel

Sun Hall Hotel

Athinon Avenue 7, 6023 Larnaca, Cyprus
Invited talk Invited

Speaker

Andreas Demou (The Cyprus Institute)

Description

Mechanistic interpretability aspires to reverse engineer AI models by breaking down black-box weights and activations into human-understandable features and circuits. A leading approach for mechanistic interpretability is sparse dictionary learning, which trains an encoder to map activations into a sparse code and a decoder to reconstruct them. The corresponding architecture is called sparse-autoencoder (SAE), and it has been shown to uncover safety-relevant concepts such as deception, bias, and harmful content, enabling targeted interventions on model behavior. In this talk, we will go through applications of SAEs in language and vision models, as well as present new SAE variants based on tensor decompositions that enhance expressivity at a fraction of the parameter cost.

Author

Andreas Demou (The Cyprus Institute)

Presentation materials

There are no materials yet.